Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Mapping tree canopy cover and canopy height with L-band SAR using LiDAR data and Random Forests

Light detection and ranging (LiDAR) data can provide direct measurements of vegetation structures but are limited by the sparse spatial coverage. Polarimetric synthetic aperture radar (SAR) can perform large-scale high-resolution mapping without weather constraints but the information about vegetation and ground subsurface are mixed in the backscatter data. In this paper, we adopted the Random Forests algorithm to train an upscaling function using tree canopy cover (TCC) and canopy height model (CHM) derived from Goddard’s LiDAR, Hyperspectral and Thermal Imager (G-LiHT) data. The regression model is then applied to the L-band Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) data acquired during the 2017 Arctic-Boreal Vulnerability Experiment (ABoVE) airborne campaign to map the TCC and CHM over the Delta Junction area in interior Alaska.

Moghaddam, Mahta↗

Validity of VO(2 max) in predicting blood volume: implications for the effect of fitness on aging

A multiple regression model was constructed to investigate the premise that blood volume (BV) could be predicted using several anthropometric variables, age, and maximal oxygen uptake (VO(2 max)). To test this hypothesis, age, calculated body surface area (height/weight composite), percent body fat (hydrostatic weight), and VO(2 max) were regressed on to BV using data obtained from 66 normal healthy men. Results from the evaluation of the full model indicated that the most parsimonious result was obtained when age and VO(2 max) were regressed on BV expressed per kilogram body weight. The full model accounted for 52% of the total variance in BV per kilogram body weight. Both age and VO(2 max) were related to BV in the positive direction. Percent body fat contributed <1% to the explained variance in BV when expressed in absolute BV (ml) or as BV per kilogram body weight. When the model was cross validated on 41 new subjects and BV per kilogram body weight was reexpressed as raw BV, the results indicated that the statistical model would be stable under cross validation (e.g., predictive applications) with an accuracy of +/- 1,200 ml at 95% confidence. Our results support the hypothesis that BV is an increasing function of aerobic fitness and to a lesser extent the age of the subject. The results may have implication as to a mechanism by which aerobic fitness and activity may be protective against reduced BV associated with aging.

NASA Discipline Cardiopulmonary↗

Gaussian-process generative model for the QCD equation of state

We develop a generative model for the nuclear matter equation of state at zero net baryon density using the Gaussian process regression method. We impose first-principles theoretical constraints from lattice quantum chromodynamics and hadron resonance gas at high- and low-temperature regions, respectively. By allowing the trained Gaussian process regression model to vary freely near the phase transition region, we generate random smooth crossover equations of state with different speeds of sound that do not rely on specific parametrizations. Here, we explore a collection of experimental observable dependencies on the generated equations of state, which paves the groundwork for future Bayesian inference studies to use experimental measurements from relativistic heavy-ion collisions to constrain the nuclear matter equation of state.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Case Study: Analysis of Autonomous Center line Tracking Neural Networks

Deep neural networks have gained widespread usage in a number of applications. However, limitations such as lack of explainability and robustness inhibit building trust in their behavior, which is crucial in safety critical applications such as autonomous driving. Therefore, techniques which aid in understanding and providing guarantees for neural network behavior are the need of the hour. In this paper, we present a case study applying a recently proposed technique, Prophecy, to analyze the behavior of a neural network model, provided by our industry partner and used for autonomous guiding of airplanes on taxi runways. This regression model takes as input an image of the runway and produces two outputs, cross-track error and heading error, which represent the position of the plane relative to the center line. We use the Prophecy tool to extract neuron activation patterns for the correctness and safety properties of the model. We show the use of these patterns to identify features of the input that explain correct and incorrect behavior. We also use the patterns to provide guarantees of consistent behavior. We explore a novel idea of using sequences of images (instead of single images) to obtain good explanations and identify regions of consistent behavior.

Deep Neural Networks↗

Investigation of the Formation of Topologically Close Packed Phase Instabilities in Nickel-Based Superalloy Rene N6

Topologically close packed (TCP) phase instability in third generation Ni-base superalloys is understood to hinder component performance when used in high-temperature jet engine applications. The detrimental effects on high temperature performance from these brittle phases includes weakening of the Ni-rich matrix through the depletion of potent solid solution strengthening elements. Thirty-four compositional variations of polycrystalline Rene N6 were defined from a design-of-experiments approach and then cast, homogenized, and finally aged to promote TCP formation. Our prior work reported on the results of the multiple retression modeling of these alloys in order to predict the volume fraction of TCP. This paper will present further regression modeling results on these alloys in order to predict the occurrence of TCP in third generation Ni-base superalloy microstructures. Kinetic results are also discussed.

Ritzert, Frank↗

Updated global and regional trends of stratospheric ozone profiles

We present updated evaluation of stratospheric ozone profile trends in the 60° S–60°N latitude range using long-term ground-based and satellite climate data records, as well as simulations by chemistry-climate models. The trends are evaluated using the LOTUS (Long-term Ozone Trends and Uncertainties in the Stratosphere) regression model. Analyses of satellite data confirm the statistically significant positive ozone trends in the period 2000–2024 in the upper stratosphere of ~1–3% per decade, with larger trends at mid-latitudes compared to the tropics. The trends are slightly positive or close to zero in the middle stratosphere, and mostly negative, -1 to -2% per decade, in the lower stratosphere, but they are not statistically significant. The morphology and magnitude of ozone trends are similar to previous analyses (2000–2020 trends). Ozone trends in 2000–2024 predicted by chemistry-climate model simulations are in good agreement with combined satellite trends. In the upper stratosphere, models predict a slightly stronger ozone recovery than observations. In the lower stratosphere, both models and satellite observations report negative trends in the tropics, while modelled ozone trends are slightly positive at mid-latitudes. Ozone profile trends over several stations estimated from ground-based records capture the same overall vertical pattern of ozone trends as merged gridded satellite datasets. Analyses of regional ozone profile trends in 2003–2024 using merged satellite datasets confirmed the previous observations of a longitudinal structure in ozone trends in the NH mid-latitude stratosphere, with positive trends over Scandinavia and negative trends over Siberia. However, the magnitude of this dipole-like structure is reduced compared to previous analyses.

trends↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

ERTS data user investigation to develop a multistage forest sampling inventory system

The author has identified the following significant results. A unique digital timber volume estimation system was developed for use with the MSS CCT tapes. The system was tested on a 64-square mile area in Northern California's Trinity Alps. The outcome of a systematic experiment, in which several possible combinations of bands 5 and 7 and a contrast measure were tried, showed that an estimated gain in precision of 50% can be obtained in a multistage sampling design. The difference between bands 5 and 7 proved to be of special importance for the estimation of biomass in the form of timber volume. In addition, an interpretation model for high flight U2 photographs was developed. A maximum multiple correlation coefficient of 0.74 was obtained for the regression model, explaining 55% of the variation in timber volume as estimated from aerial photos and ground measurments. An interpretation model for MSS color composites is in the testing stage.

Langley, P. G.↗

Understanding and Verifying Neural Networks

Deep Neural Networks (DNNs) have gained immense popularity in recent times and have widespread use in applications such as image classification, sentiment analysis, speech recognition and also in safety-critical applications such as autonomous driving. However, they suffer limitations such as lack of explainability and robustness which raise safety and security concerns in their usage. Further, the complex structure and large input spaces of DNNs act as an impediment to thorough verification and testing. The SafeDNN project at the Robust Software Engineering (RSE) group at NASA aims at exploring techniques to ensure that systems that use deep neural networks are safe, robust and interpretable. In this talk, I will be presenting our technique Prophecy that automatically infers formal properties of deep neural network models. The tool extracts patterns based on neuron activations as preconditions that imply certain desirable output properties of the model. I would be highlighting case studies that use Prophecy in obtaining explanations for network decisions, understanding correct and incorrect behavior, providing formal guarantees wrt safety and robustness, and debugging neural network models. We have applied the tool on image classification networks, neural network controllers providing turn advisories in unmanned aircrafts, regression models used for autonomous center-line tracking in aircrafts and neural network object detectors

Deep Neural Networks↗

Regression Analysis of Long-Term Profile Ozone Data Set from BUV Instruments

We have produced a profile merged ozone data set (MOD) based on the SBUV/SBUV2 series of nadir-viewing satellite backscatter instruments, covering the period from November 1978 - December 2003. In 2004, data from the Nimbus 7 SBUV and NOAA 9, ll, and 16 SBUV/2 instruments were reprocessed using the Version 8 (V8) algorithm and most recent calibrations. More recently, data from the Nimbus 4 BUT instrument, which was operational from 1970 - 1977, were also reprocessed using the V8 algorithm. As part of the V8 profile calibration, the Nimbus 7 and NOAA 9 (1993-1997 only) instrument calibrations have been adjusted to match the NOAA 11 calibration, which was established based on comparisons with SSBUV shuttle flight data. Differences between NOAA 11, Nimbus 7 and NOAA 9 profile zonal means are within plus or minus 5% at all levels when averaged over the respective periods of data overlap. NOAA 16 SBUV/2 data have insufficient overlap with NOAA 11, so its calibration is based on pre-flight information. Mean differences over 4 months of overlap are within plus or minus 7%. Given the level of agreement between the data sets, we simply average the ozone values during periods of instrument overlap to produce the MOD profile data set. Initial comparisons of coincident matches of N4 BUV and Arosa Umkehr data show mean differences of 0.5 (0.5)% at 30km; 7.5 (0.5)% at 35 km; and 11 (0.7)% at 40 km, where the number in parentheses is the standard error of the mean. In this study, we use the MOD profile data set (1978-2003) to estimate the change in profile ozone due to changing stratospheric chlorine levels. We use a standard linear regression model with proxies for the seasonal cycle, solar cycle, QBO, and ozone trend. To account for the non-linearity of stratospheric chlorine levels since the late 1990s, we use a time series of Effective Chlorine, defined as the global average of Chlorine + 50 * Bromine at 1 hPa, as the trend proxy. The Effective Chlorine data are taken from the 3-D Goddard CTM. We will show the latest trend results using this statistical model. In addition, the Nimbus 4 BUV data offer an opportunity to test the physical properties of our statistical model. From ground-based comparisons we will establish an uncertainty range for the Nimbus 4 data. We then extrapolate our statistical model fit backwards in time and compare to the Nimbus 4 data. We compare the characteristics of the residual, defined as the difference between the data and statistical regression fit, during the Nimbus 4 time period and the 1978-2003 period over which the statistical model coefficients were estimated, and present these results.

Stolarski, Richard S.↗

The northeast materials database for magnetic materials

The discovery of magnetic materials with high operating temperature ranges and optimized performance is essential for advanced applications. Current data-driven approaches are limited by the lack of accurate, comprehensive, and feature-rich databases. This study aims to address this challenge by using Large Language Models (LLMs) to create a comprehensive, experiment-based, magnetic materials database named the Northeast Materials Database (NEMAD), which consists of 67,573 magnetic materials entries (www.nemad.org). The database incorporates chemical composition, magnetic phase transition temperatures, structural details, and magnetic properties. Enabled by NEMAD, we trained machine learning models to classify materials and predict transition temperatures. Our classification model achieved an accuracy of 90% in categorizing materials as ferromagnetic (FM), antiferromagnetic (AFM), and non-magnetic (NM). The regression models predict Curie (Néel) temperature with a coefficient of determination (R 2 ) of 0.87 (0.83) and a mean absolute error (MAE) of 56K (38K). These models identified 25 (13) FM (AFM) candidates with a predicted Curie (Néel) temperature above 500K (100K) from the Materials Project. This work shows the feasibility of combining LLMs for automated data extraction and machine learning models to accelerate the discovery of magnetic materials.

Ferromagnetism↗

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES↗

Analysis of upper stratospheric Umkehr ozone profile data for trends and the effects of stratospheric aerosols

The effect of stratospheric aerosols on Umkehr estimates of long-term ozone depletion associated with chlorofluoromethanes (CFMs) is considered in a statistical time series trend analysis. Time series models are estimated using monthly averages of Umkehr measurements made over the last 15 to 20 years. The time series regression models incorporate seasonal, trend and noise factors and an additional factor to account for the effects of atmospheric aerosols on the Umkehr measurements. The analysis indicates a statistically significant relation with atmospheric aerosol transmission in the Umkehr layers and implies that the relation is an important factor in any time series trend analysis of Umkehr data. Taking this relation into account, statistically significant negative trends were found in the Upper Umkehr layers. It is pointed out that upper stratospheric ozone could be sensitive to long-term solar variability as well as other possible influences in addition to CFM-induced effects, and therefore the cause or causes of the estimated trend cannot be unambiguously estimated using current Umkehr data.

Reinsel, G. C.↗

The Role of Hierarchy in Response Surface Modeling of Wind Tunnel Data

This paper is intended as a tutorial introduction to certain aspects of response surface modeling, for the experimentalist who has started to explore these methods as a means of improving productivity and quality in wind tunnel testing and other aerospace applications. A brief review of the productivity advantages of response surface modeling in aerospace research is followed by a description of the advantages of a common coding scheme that scales and centers independent variables. The benefits of model term reduction are reviewed. A constraint on model term reduction with coded factors is described in some detail, which requires such models to be well-formulated, or hierarchical. Examples illustrate the consequences of ignoring this constraint. The implication for automated regression model reduction procedures is discussed, and some opinions formed from the author s experience are offered on coding, model reduction, and hierarchy.

DeLoach, Richard↗

Artificial intelligence-based predictive modeling for imaging neutral particle analyzers on the DIII-D tokamak

The Imaging Neutral Particle Analyzer (INPA) at DIII-D is a diagnostic system used to accurately resolve the energy and spatial distributions of fast ions in fusion plasmas. A novel artificial intelligence (AI) technique named INPA-net is based on Reservoir Computing Networks and developed here to predict active and passive signals produced by charge-exchange reactions from injected and edge-cold neutrals, respectively, in magnetically confined fusion plasmas. This model is trained using a set of 21 time domain signals between 0 s to 3.35 s that includes injected beam and thermal plasma information, and 6444 real 2D experimental images of the INPA in 12 plasma discharges at DIII-D. The trained neural network is able to forecast experimental images in real-time. The model achieves an R-squared value of 0.91, which is higher than the 0.83 value achieved by a simple linear regression model. This improvement highlights the model's enhanced predictive accuracy for measured images from the validation set. This AI approach is valuable due to its rapid response times and potential for integration into real-time plasma control systems. A version of this model capable of generating syntehic images would be useful for the real-time monitoring of fast-ion transport. A comprehensive sensitivity study reveals that INPA-net maintains high performance even with variations in the input parameters, indicating the model's robustness and reliability. While developed for the INPA, the underlying architecture is adaptable and may be applied to various 2D imaging diagnostics in fusion research.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The Prediction Properties of Inverse and Reverse Regression for the Simple Linear Calibration Problem

The calibration of measurement systems is a fundamental but under-studied problem within industrial statistics. The origins of this problem go back to basic chemical analysis based on NIST standards. In today's world these issues extend to mechanical, electrical, and materials engineering. Often, these new scenarios do not provide "gold standards" such as the standard weights provided by NIST. This paper considers the classic "forward regression followed by inverse regression" approach. In this approach the initial experiment treats the "standards" as the regressor and the observed values as the response to calibrate the instrument. The analyst then must invert the resulting regression model in order to use the instrument to make actual measurements in practice. This paper compares this classical approach to "reverse regression," which treats the standards as the response and the observed measurements as the regressor in the calibration experiment. Such an approach is intuitively appealing because it avoids the need for the inverse regression. However, it also violates some of the basic regression assumptions.

Parker, Peter A.↗

Temperature-dependent mechanical properties and crystal plasticity parameters for additively manufactured Haynes-214 alloy: Experiments and numerical modeling

Our experimental mechanical testing data demonstrated that the additively manufactured (AM) laser powder bed fusion (L-PBF) Haynes-214 alloy exhibits non-linear mechanical properties as the temperature rises from ambient to 870 °C. Crystal plasticity (CP) simulations provide an effective approach to gaining deeper insights into microstructure-property linkages under thermomechanical loading. This method can reduce the need for costly high-temperature mechanical testing while accounting for the effects of crystallographic texture and grain morphology on the mechanical behavior of AM materials. However, calibrating a CP model is time-consuming because individual simulations are computationally expensive and hundreds (or more) of iterations over parameter sets may be required. To address this issue, we have designed a machine learning-differential evolution (ML-DE) CP framework that can accurately interpolate the tensile properties of AM L-PBF Haynes-214 alloy across a wide temperature range from ambient to 870 °C, with minimal reliance on experimental data. The framework uses electron backscatter diffraction (EBSD) measurements to generate statistically equivalent microstructural volume elements to serve as inputs to the CP modeling framework. Stress–strain curves were generated from 1000 CP simulations, which serve as the training data set for the three ML regression algorithms explored: linear, extra-trees, and multi-layer perceptron. These three regression models were independently evaluated to compare their efficiency and identify the most suitable algorithm for the given problem. Results revealed that the extra-trees ML regressor outperforms the other models in both qualitative and quantitative aspects with an R 2 of 0.98. Subsequently, the differential evolution optimization approach is employed to calibrate the ML-based CP material parameters with experimental results obtained at various temperatures. Finally, temperature-dependent CP material parameters are formulated. The effectiveness and efficiency of the designed framework are validated through comparison with experimental results, demonstrating a high degree of agreement. These calibrated parametric constitutive equations enable further use of the CP model to study the deformation behavior of this alloy under a wide range of thermo-mechanical loading conditions.

36 MATERIALS SCIENCE↗

Rapid monitoring of fermentations: a feasibility study on biological 2,3-butanediol production

2,3-butanediol (2,3-BDO) is an economically important platform chemical that can be produced by the fermentation of sugars using an engineered strain of Zymomonas mobilis . These fermentations require continuous monitoring and modification of fermentation conditions to maximize 2,3-BDO yields and minimize the production of the undesired coproducts glycerol and acetoin. Because of the time required for sampling and off-line chromatographic measurement of fermentation samples, the ability of fermentation scientists to modify fermentation conditions in a timely manner is limited. The goal of this study was to test if near-infrared spectroscopy (NIRS) along with multivariate statistics could reduce the time needed for this analysis and enable real-time monitoring and control of the fermentation. In this work we developed partial least squares (PLS) calibration models to predict the concentrations of glucose, xylose, 2,3-BDO, acetoin, and glycerol in fermentations via NIRS using two different spectrometers and two different spectroscopy modalities. We first evaluated the feasibility of rapid NIRS monitoring through experiments where we measured the signals from each analyte of interest and built NIRS-based PLS models using spectra from synthetic samples containing uncorrelated concentrations of these analytes. All analytes showed unique spectral signatures, and this initial modeling showed that all analytes could be detected simultaneously. We then began work with samples from laboratory fermentation experiments and tested the feasibility of regression model development across two spectral collection modalities (at-line and on-line) and two instruments: a laboratory-grade instrument and a low-cost instrument with a more limited spectral range. All modalities showed promise in the ability to monitor Z. mobilis fermentations of glucose and xylose to 2,3-BDO. The low-cost instrument displayed a lower signal-to-noise ratio than the laboratory-grade instrument, which led to comparatively lower performance overall, but still provided sufficient accuracy to monitor fermentation trends. While the ease of use of on-line monitoring systems was favored as compared to at-line systems due to the lack of sampling required and potential for automated process control, we observed some decrease in performance due to the additional complexity of the sample matrix. We have demonstrated that NIRS combined with multivariate analysis can be used for at-line and on-line monitoring of the concentrations of glucose, xylose, 2,3-BDO, acetoin, and glycerol during Z. mobilis fermentations. The decrease in signal-to-noise ratio when using a low-cost spectrometer led to greater prediction error than the laboratory-grade spectrometer for at-line monitoring. The on-line monitoring modality showed great promise for real time process control via NIRS.

09 BIOMASS FUELS↗