Search NASA⌕ Search

SEARCH · Search NASA

Results for “regression analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40

Acousto-ultrasonic verification of the strength of filament wound composite material

The concept of acousto-ultrasonic (AU) waveform partitioning was applied to nondestructive evaluation of mechanical properties in filament wound composites (FWC). A series of FWC test specimens were subjected to AU analysis and the results were compared with destructively measured interlaminar shear strengths (ISS). AU stress-wave factor (SWF) measurements gave greater than 90 percent correlation coefficient upon regression against the ISS. This high correlation was achieved by employing the appropriate time and frequency domain partitioning as dictated by wave propagation path analysis. There is indication that different SWF frequency partitions are sensitive to ISS at different depths below the surface.

Kautz, H. E.↗

Experiments with Test Case Generation and Runtime Analysis

Software testing is typically an ad hoc process where human testers manually write many test inputs and expected test results, perhaps automating their execution in a regression suite. This process is cumbersome and costly. This paper reports preliminary results on an approach to further automate this process. The approach consists of combining automated test case generation based on systematically exploring the program's input domain, with runtime analysis, where execution traces are monitored and verified against temporal logic specifications, or analyzed using advanced algorithms for detecting concurrency errors such as data races and deadlocks. The approach suggests to generate specifications dynamically per input instance rather than statically once-and-for-all. The paper describes experiments with variants of this approach in the context of two examples, a planetary rover controller and a space craft fault protection system.

Artho, Cyrille↗

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Meteoric 10Be Flux Calibration Data for the East River Watershed, Colorado, USA

This data package contains tabular and geospatial data used to quantify and model meteoric beryllium-10 fluxes in the East River watershed, Colorado, USA. The tabular component includes calibration-site data from five glacial moraine sites and includes environmental variables used to evaluate spatial controls on meteoric 10Be delivery, including elevation, mean annual precipitation (MAP), mean snow depth, and mean snow water equivalent (SWE). These site-level data were used to compare observed fluxes with environmental gradients across the watershed and to evaluate the effects of erosion correction on flux estimates. The package also includes supporting slope and curvature values used to assess topographic inputs to the erosion analysis. A second component of the data package contains updated manuscript tables and regression outputs used to summarize the relationships between meteoric 10Be flux and environmental predictors. These tables include meteoric 10Be sample information and AMS results, site-level environmental values, site-level meteoric 10Be inventory and flux values, watershed-averaged predicted fluxes, soil bulk density measurements, fine-fraction values, soil pH measurements, and regression statistics including slope, intercept, coefficient of determination, and p-value. The regression products include both standard linear regressions and regressions in which the intercept is constrained to pass through zero, and they support the analyses presented in the companion manuscript. Together, these tabular files provide the numerical basis for the manuscript tables and the regression-based interpretation of meteoric 10Be flux variability in a snow-dominated mountain watershed. The geospatial component of the package consists of GeoTIFF raster files used to generate the map products presented in Figures 2 and 6 of the companion manuscript. These rasters represent watershed-scale spatial layers for environmental variables and regression-based predictions of meteoric 10Be flux. This dataset contains comma-separated values files (.csv), Microsoft Excel files (.xlsx), GeoTIFF raster files (.tif), and upporting metadata files, including CSV data dictionaries and readme text files (.csv, .txt). The tabular files can be opened with standard spreadsheet software, and the raster files can be viewed and analyzed in GIS software such as ArcGIS Pro or QGIS. Together, these files document the numerical and spatial datasets used to calibrate and predict meteoric 10Be delivery in the East River watershed.

East River↗

Refining Fast Simulation Using Machine Learning

A growing reliance on the fast Monte Carlo (FastSim) will accompany the high luminosity and detector granularity expected in Phase 2. FastSim is roughly 10 times faster than equivalent GEANT4-based full simulation (FullSim). However, reduced accuracy of the FastSim affects some analysis variables and collections. To improve its accuracy, FastSim is refined using regression-based neural networks trained with ML. The status of FastSim refinement is presented. The results show improved agreement with the FullSim output and an improvement in correlations among output observables and external parameters.

Güngördü, Acelya Deniz↗

Mining Product Reviews for Important Product Features of Refurbished iPhones

Problem: Remanufacturers want to increase consumer interest in refurbished products, which motivates the need to understand which product features are important to buyers of refurbished products such as mobile phones. Research Questions: This study addresses two questions. First, which product features are most important for buyers of refurbished iPhones? Second, how do those preferences differ from the preferences of buyers of new iPhones? Methods: Online reviews of iPhones are obtained and converted into a document–term matrix. Using this text model, three subsets of features are identified using statistical analysis of frequency of mention: most frequent, average, and least frequent. A logistic regression (LR) model is then used to identify which features are most predictive of whether a review is for a new or refurbished phone. Results: Buyers of refurbished phones mention battery health, screen/display, shell condition, and brand significantly more often than other features. Directly contrasting reviews of refurbished versus new phones shows that shell condition, brand, speaker, and charger are found to be the most predictive product features indicated in reviews for refurbished phones. Of those, the shell condition is significantly more predictive than the others. Implications: The results identify product features that remanufacturers of iPhones can emphasize to increase customer demand.

Anisi, Atefeh↗

Research in the application of spectral data to crop identification and assessment, volume 2

The development of spectrometry crop development stage models is discussed with emphasis on models for corn and soybeans. One photothermal and four thermal meteorological models are evaluated. Spectral data were investigated as a source of information for crop yield models. Intercepted solar radiation and soil productivity are identified as factors related to yield which can be estimated from spectral data. Several techniques for machine classification of remotely sensed data for crop inventory were evaluated. Early season estimation, training procedures, the relationship of scene characteristics to classification performance, and full frame classification methods were studied. The optimal level for combining area and yield estimates of corn and soybeans is assessed utilizing current technology: digital analysis of LANDSAT MSS data on sample segments to provide area estimates and regression models to provide yield estimates.

Daughtry, C. S. T.↗

Effects of alloy composition on cyclic flame hot-corrosion attack of cast nickel-base superalloys at 900 deg C

The effects of Cr, Al, Ti, Mo, Ta, Nb, and W content on the hot corrosion of nickel base alloys were investigated. The alloys were tested in a Mach 0.3 flame with 0.5 ppmw sodium at a temperature of 900 C. One nondestructive and three destructive tests were conducted. The best corrosion resistance was achieved when the Cr content was 12 wt %. However, some lower-Cr-content alloys ( 10 wt%) exhibited reasonable resistance provided that the Al content alloys ( 10 wt %) exhibited reasonable resistance provided that the Al content was 2.5 wt % and the Ti content was Aa wt %. The effect of W, Ta, Mo, and Nb contents on the hot-corrosion resistance varied depending on the Al and Ti contents. Several commercial alloy compositions were also tested and the corrosion attack was measured. Predicted attack was calculated for these alloys from derived regression equations and was in reasonable agreement with that experimentally measured. The regression equations were derived from measurements made on alloys in a one-quarter replicate of a 2(7) statistical design alloy composition experiment. These regression equations represent a simple linear model and are only a very preliminary analysis of the data needed to provide insights into the experimental method.

Deadmore, D. L.↗

A statistical study of the surface accuracy of a planar truss beam

Surface error statistics for single-layer and double-layer planar truss beams with random member-length errors were calculated using a Monte-Carlo technique in conjunction with finite-element analysis. Surface error was calculated in terms of the normal distance from a regression line to the surface nodes of the distorted beam. Results for both single-layer and double-layer beams indicate that a minimum root-mean-square surface error can be achieved by optimizing the depth-to-length ratio of a truss beam. The statically indeterminate double-layer beams can provide greater surface accuracy, though at the expense of significantly greater complexity.

Kenner, W. Scott↗

Analysis of a Split-Plot Experimental Design Applied to a Low-Speed Wind Tunnel Investigation

A procedure to analyze a split-plot experimental design featuring two input factors, two levels of randomization, and two error structures in a low-speed wind tunnel investigation of a small-scale model of a fighter airplane configuration is described in this report. Standard commercially-available statistical software was used to analyze the test results obtained in a randomization-restricted environment often encountered in wind tunnel testing. The input factors were differential horizontal stabilizer incidence and the angle of attack. The response variables were the aerodynamic coefficients of lift, drag, and pitching moment. Using split-plot terminology, the whole plot, or difficult-to-change, factor was the differential horizontal stabilizer incidence, and the subplot, or easy-to-change, factor was the angle of attack. The whole plot and subplot factors were both tested at three levels. Degrees of freedom for the whole plot error were provided by replication in the form of three blocks, or replicates, which were intended to simulate three consecutive days of wind tunnel facility operation. The analysis was conducted in three stages, which yielded the estimated mean squares, multiple regression function coefficients, and corresponding tests of significance for all individual terms at the whole plot and subplot levels for the three aerodynamic response variables. The estimated regression functions included main effects and two-factor interaction for the lift coefficient, main effects, two-factor interaction, and quadratic effects for the drag coefficient, and only main effects for the pitching moment coefficient.

Erickson, Gary E.↗

On the Feasibility of Monitoring Carbon Monoxide in the Lower Troposphere from a Constellation of Northern Hemisphere Geostationary Satellites (PART 1)

By the end of the current decade, there are plans to deploy several geostationary Earth orbit (GEO) satellite missions for atmospheric composition over North America, East Asia and Europe with additional missions proposed. Together, these present the possibility of a constellation of geostationary platforms to achieve continuous time-resolved high-density observations over continental domains for mapping pollutant sources and variability at diurnal and local scales. In this paper, we use a novel approach to sample a very high global resolution model (GEOS-5 at 7 km horizontal resolution) to produce a dataset of synthetic carbon monoxide pollution observations representative of those potentially obtainable from a GEO satellite constellation with predicted measurement sensitivities based on current remote sensing capabilities. Part 1 of this study focuses on the production of simulated synthetic measurements for air quality OSSEs (Observing System Simulation Experiments). We simulate carbon monoxide nadir retrievals using a technique that provides realistic measurements with very low computational cost. We discuss the sampling methodology: the projection of footprints and areas of regard for geostationary geometries over each of the North America, East Asia and Europe regions; the regression method to simulate measurement sensitivity; and the measurement error simulation. A detailed analysis of the simulated observation sensitivity is performed, and limitations of the method are discussed. We also describe impacts from clouds, showing that the efficiency of an instrument making atmospheric composition measurements on a geostationary platform is dependent on the dominant weather regime over a given region and the pixel size resolution. These results demonstrate the viability of the "instrument simulator" step for an OSSE to assess the performance of a constellation of geostationary satellites for air quality measurements.

GEOS-5↗

Shergottite Lead Isotope Signature in Chassigny and the Nakhlites

The nakhlites/chassignites and the shergottites represent two differing suites of basaltic martian meteorites. The shergottites have ages less than or equal to 0.6 Ga and a large range of initial Sr-/Sr-86 and epsilon (Nd-143) ratios. Conversely, the nakhlites and chassignites cluster at 1.3-1.4 Ga and have a limited range of initial Sr-87/Sr-86 and epsilon (Nd-143). More importantly, the shergottites have epsilon (W-182) less than 1, whereas the nakhlites and chassignites have epsilon (W-182) approximately 3. This latter observation precludes the extraction of both meteorite groups from a single source region. However, recent Pb isotopic analyses indicate that there may have been interaction between shergottite and nakhlite/chassignite Pb reservoirs.Pb Analyses of Chassigny: Two different studies haveinvestigated 207Pb/204Pb vs. 206Pb/204Pb in Chassigny: (i)TIMS bulk-rock analyses of successive leaches and theirresidue [3]; and (ii) SIMS analysis of individual minerals[4]. The bulk-rock analyses fall along a regression of SIMSplagioclase analyses that define an errorchron that is olderthan the Solar System (4.61±0.1 Ga); i.e., these define amixing line between Chassigny’s principal Pb isotopic components(Fig. 1). Augites and olivines in Chassingy (notshown) also fall along or near the plagioclase regression [4].This agreement indicates that the whole-rock leachateslikely measure indigenous, martian Pb, not terrestrial contamination[5]. SIMS analyses of K-spars and sulfides definea separate, sub-parallel trend having higher 207Pb/206Pbvalues ([4]; Fig. 1). The good agreement between the bulkrockanalyses and the SIMS analyses of plagioclases alsoindicates that the Pb in the K-spars and sulfides cannot be amajor component of Chassigny.The depleted reservoir sampled by Chassigny plagioclaseis not the same as the solar system initial (PAT) andrequires a multi-stage origin. Here we show a two-stagemodel (Fig. 1) with a 238U/204Pb (μ) of 0.5 for 4.5-2.4 Gaand a μ of 7 for 2.4-1.4 Ga. This is not a unique model butdoes produce a Pb composition that falls on the plagioclaseregression at 1.4 Ga, the approximate igneous age of Chassigny [1]. It should be noted that low-μ single-stage modelsare not capable of producing sufficiently radiogenic 206Pb/204Pb at 1.4 Ga.Relation to Shergottites: The Chassigny K-spars and sulfides fall along a second mixing line defined by leachesand residues of depleted and intermediate shergottites [6]. This mixing line falls above the plagioclase regression.Therefore, we also interpret the radiogenic component of this mixing line to represent indigenous martian Pb. It ispossible that the depleted and intermediate shergottites and the Chassigny plagioclases sample radiogenic Pb from thethe same source, i.e., the mixing lines may intersect at high 206Pb/204Pb.Both K-spar and sulfide are late-stage phases. At the time of their crystallization, the Chassigny system appearsto have remained open to a depleted shergottite Pb reservoir. The depleted component of the shergottite mixing linecan be generated by a single-stage evolution from PAT (4.5 to 1.4 Ga) in a reservoir having a μ ~2. A similar modelfor the most depleted shergottites is also possible: μ = 1.5 for 4.5 to 0.3 Ga.Nakhlites: Nakhlite analyses plot between the shergottite and Chassigny plagioclase regressions [3]. So again,members of the nakhlite/chassignite suite show affinities to shergottite Pb.

Jones, J. H.↗

Evaluating Approaches Relating Ecosystem Productivity with Desis Spectral Information

Data from the DLR Earth Sensing Imaging Spectrometer (DESIS), mounted on the International Space Station (ISS), were used to develop and test algorithms for remotely retrieving ecosystem productivity. Twenty DESIS images were used from three widely separated forested study sites representing deciduous and conifer forests. Gross primary production (GPP) values from eddy covariance flux towers at the sites were matched with DESIS spectral reflectances collected on the same days. Multiple algorithms were successful relating spectral reflectance with GPP, including: spectral vegetation indices (SVI) sensitive to chlorophyll content, SVI used in a photosynthetic light-use efficiency model framework, spectral shape characteristics through spectral derivatives and absorption feature analysis, and statistical models leading to multiband hyperspectral indices from partial least squares regression. Successful algorithms were able to achieve R2 better than 0.7 using a diverse set of observations combining data from different sites from multiple years and at multiple times during the year. The demonstrated robustness of the algorithms provides some confidence in using DESIS imagery to map spatial patterns of GPP.

K F Huemmrich↗

Evaluating Approaches Relating Ecosystem Productivity with DESIS Spectral Information

Data from the DLR Earth Sensing Imaging Spectrometer (DESIS), mounted on the International Space Station (ISS), were used to develop and test algorithms for remotely retrieving ecosystem productivity. Twenty DESIS images were used from three widely separated forested study sites representing deciduous and conifer forests. Gross primary production (GPP) values from eddy covariance flux towers at the sites were matched with DESIS spectral reflectances collected on the same days. Multiple algorithms were successful relating spectral reflectance with GPP, including: spectral vegetation indices (SVI) sensitive to chlorophyll content, SVI used in a photosynthetic light-use efficiency model framework, spectral shape characteristics through spectral derivatives and absorption feature analysis, and statistical models leading to multiband hyperspectral indices from partial least squares regression. Successful algorithms were able to achieve R2 better than 0.7 using a diverse set of observations combining data from different sites from multiple years and at multiple times during the year. The demonstrated robustness of the algorithms provides some confidence in using DESIS imagery to map spatial patterns of GPP.

Gross Primary Productivity (GPP)↗

Visualization and Quantification of Wind Induced Variability in Hydrogen Clouds Following Releases of Liquid Hydrogen: Preprint

Well characterized experimental data for consequence model validation is important in progressing the use of liquid hydrogen as an energy carrier. In 2019, the Health and Safety Executive (HSE) undertook a series of liquid hydrogen dispersion and combustion experiments as a part of the Pre-normative Research into the Safe Use of Liquid Hydrogen (PRESLHY) project. In partnership between the National Renewable Energy Laboratory (NREL) and HSE, time and spatially varying hydrogen concentration measurements were made in 25 dispersion experiments and 23 congested ignition experiments associated with PRESLHY WP3 and WP5, respectively. These measurements were undertaken using the hydrogen wide area monitoring system developed by NREL. During the 23 congested ignition experiments, high variability was observed in the measured explosion severity during experiments with similar initial conditions. This led to the conclusion that wind, including localized gusts, had a large influence on the dispersion of the hydrogen, and therefore the quantity of hydrogen that was present in the congested region of the explosions. Using the hydrogen concentration measurements taken immediately prior to ignition, the hydrogen clouds were visualized in an attempt to rationalize the variability in overpressure between the tests. Gaussian process regression was applied to quantify the variability of the measured hydrogen concentrations. This analysis could also be used to guide modifications in experimental designs for future research on hydrogen combustion behavior.

HSR&D↗

A Provably Accurate Randomized Sampling Algorithm for Logistic Regression

In statistics and machine learning, logistic regression is a widely-used supervised learning technique primarily employed for binary classification tasks. When the number of observations greatly exceeds the number of predictor variables, we present a simple, randomized sampling-based algorithm for logistic regression problem that guarantees high-quality approximations to both the estimated probabilities and the overall discrepancy of the model. Our analysis builds upon two simple structural conditions that boil down to randomized matrix multiplication, a fundamental and well-understood primitive of randomized numerical linear algebra. We analyze the properties of estimated probabilities of logistic regression when leverage scores are used to sample observations, and prove that accurate approximations can be achieved with a sample whose size is much smaller than the total number of observations. To further validate our theoretical findings, we conduct comprehensive empirical evaluations. Overall, our work sheds light on the potential of using randomized sampling approaches to efficiently approximate the estimated probabilities in logistic regression, offering a practical and computationally efficient solution for large-scale datasets.

Chowdhury, Agniva↗

Optimization and Multimachine Learning Algorithms to Predict Nanometal Surface Area Transfer Parameters for Gold and Silver Nanoparticles

Interactions between gold metallic nanoparticles and molecular dyes have been well described by the nanometal surface energy transfer (NSET) mechanism. However, the expansion and testing of this model for nanoparticles of different metal composition is needed to develop a greater variety of nanosensors for medical and commercial applications. In this study, the NSET formula was slightly modified in the size-dependent dampening constant and skin depth terms to allow for modeling of different metals as well as testing the quenching effects created by variously sized gold, silver, copper, and platinum nanoparticles. Overall, the metal nanoparticles followed more closely the NSET prediction than for Förster resonance energy transfer, though scattering effects began to occur at 20 nm in the nanoparticle diameter. To further improve the NSET theoretical equation, an attempt was made to set a best-fit line of the NSET theoretical equation curve onto the Au and Ag data points. An exhaustive grid search optimizer was applied in the ranges for two variables, 0.1≤C≤2.0 and 0≤α≤4, representing the metal dampening constant and the orientation of donor to the metal surface, respectively. Three different grid searches, starting from coarse (entire range) to finer (narrower range), resulted in more than one million total calculations with values C=2.0 and α=0.0736. The results improved the calculation, but further analysis needed to be conducted in order to find any additional missing physics. With that motivation, two artificial intelligence/machine learning (AI/ML) algorithms, multilayer perception and least absolute shrinkage and selection operator regression, gave a correlation coefficient, R2, greater than 0.97, indicating that the small dataset was not overfitting and was method-independent. This analysis indicates that an investigation is warranted to focus on deeper physics informed machine learning for the NSET equations.

Demers, Steven M. E. (ORCID:0000000192213246)↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗