Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Decayheatml

This code is designed to predict and analyze the decay heat generated in molten salt reactors (MSRs) using a hybrid approach that combines machine learning and segmented polynomial fitting. The accurate prediction of decay heat is essential for reactor safety and the optimization of spent fuel storage. The code operates through several key components: 1) Data Architecture: It incorporates a modular data architecture that handles various MSR-specific operational parameters such as power density, humidity content, and air ingress. These parameters are sampled using Sobol sequences to ensure comprehensive coverage of operational uncertainties. 2) Machine Learning Framework: The code employs a diverse set of machine learning models, including polynomial regression, decision trees, random forests, gradient boosting, support vector regression, k-nearest neighbors, multi-layer perceptrons, and symbolic regression. These models are trained to predict decay heat over a wide temporal range, from immediate shutdown up to 10,000 years. 3) Region-Optimized Training: The temporal domain is divided into multiple regions, each modeled separately to capture distinct decay heat characteristics across different time scales. This approach significantly improves the accuracy and interpretability of predictions. 4) Segmented Polynomial Interpretation (SPI): The SPI method translates machine learning predictions into piecewise polynomial equations. These equations are physically interpretable and can be directly integrated into existing engineering workflows and safety analyses. 5) Front-End Interfaces: The code includes both a Jupyter notebook interface for research development and a Streamlit web application for operational deployment. These interfaces allow users to interactively explore decay heat predictions, adjust operational parameters, and visualize results in real-time. 6) Applications: The framework supports various applications, including safety system validation and spent fuel container optimization. It enables real-time evaluation of worst-case decay heat scenarios, informing the design of passive safety systems and optimizing container designs for long-term storage. Overall, this code provides a robust, accurate, and user-friendly tool for predicting decay heat in MSRs, enhancing reactor safety, and optimizing spent fuel management.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

Modeling the Height of Young Forests Regenerating from Recent Disturbances in Mississippi using Landsat and ICESat data

Many forestry and earth science applications require spatially detailed forest height data sets. Among the various remote sensing technologies, lidar offers the most potential for obtaining reliable height measurement. However, existing and planned spaceborne lidar systems do not have the capability to produce spatially contiguous, fine resolution forest height maps over large areas. This paper describes a Landsat-lidar fusion approach for modeling the height of young forests by integrating historical Landsat observations with lidar data acquired by the Geoscience Laser Altimeter System (GLAS) instrument onboard the Ice, Cloud, and land Elevation (ICESat) satellite. In this approach, "young" forests refer to forests reestablished following recent disturbances mapped using Landsat time-series stacks (LTSS) and a vegetation change tracker (VCT) algorithm. The GLAS lidar data is used to retrieve forest height at sample locations represented by the footprints of the lidar data. These samples are used to establish relationships between lidar-based forest height measurements and LTSS-VCT disturbance products. The height of "young" forest is then mapped based on the derived relationships and the LTSS-VCT disturbance products. This approach was developed and tested over the state of Mississippi. Of the various models evaluated, a regression tree model predicting forest height from age since disturbance and three cumulative indices produced by the LTSS-VCT method yielded the lowest cross validation error. The R(exp 2) and root mean square difference (RMSD) between predicted and GLAS-based height measurements were 0.91 and 1.97 m, respectively. Predictions of this model had much higher errors than indicated by cross validation analysis when evaluated using field plot data collected through the Forest Inventory and Analysis Program of USDA Forest Service. Much of these errors were due to a lack of separation between stand clearing and non-stand clearing disturbances in current LTSS-VCT products and difficulty in deriving reliable forest height measurements using GLAS samples when terrain relief was present within their footprints. In addition, a systematic underestimation of about 5 m by the developed model was also observed, half of which could be explained by forest growth that occurred between field measurement year and model target year. The remaining difference suggests that tree height measurements derived using waveform lidar data could be significantly underestimated, especially for young pine forests. Options for improving the height modeling approach developed in this study were discussed.

Li, Ainong↗

A Data-Driven Method for Modeling Creep-Fatigue Stress- Strain Behavior Using Neural ODEs

In this paper, we introduce a data-driven machine learning approach for modeling one-dimensional stress–strain behavior under cyclic loading, utilizing experimental data from the nickel-based Alloy 617. The study employs uniaxial creep–fatigue test data acquired under various loading histories and compares two distinct neural network-based ODE models. The first model, known as the black-box model, comprehensively describes the strain–stress relationship using a Neural ODE equation. To interpret this black-box model, we apply the Sparse Identification of Nonlinear Dynamical Systems (SINDy) technique, transforming the black-box model into an equation-based model using symbolic regression. The second model, the Neural flow rule model, incorporates Hooke’s Law for the linear elastic component, with the nonlinear part characterized by a Neural ODE. Both models are trained with experimental data to accurately reflect the observed stress–strain behavior. We conduct a detailed comparison with the standard Chaboche model, which includes three back stresses. Our results demonstrate that the neural network-based ODE models precisely capture the experimental creep–fatigue mechanical behavior, exceeding the standard Chaboche model’s accuracy. Furthermore, an interpretable model derived from the black-box neural ODE model through symbolic regression achieves accuracy comparable to the Chaboche model, enhancing its interpretability. The results highlight the potential of neural network-based ODE models to depict complex creep–fatigue behavior, eliminating the necessity for experts to define a specific, material-focused model form.

creep-fatigue↗

Efficient data-driven regression for reduced-order modeling of spatial pattern formation

We present an efficient data-driven regression approach for constructing reduced-order models (ROMs) of reaction-diffusion systems exhibiting pattern formation. The ROMs are learned non-intrusively from available training data of physically accurate numerical simulations. The method can be applied to general nonlinear systems through the use of polynomial model form, while not requiring knowledge of the underlying physical model, governing equations, or numerical solvers. The process of learning ROMs is posed as a low-cost least-squares problem in a reduced-order subspace identified via Proper Orthogonal Decomposition (POD). Numerical experiments on classical pattern-forming systems–including the Schnakenberg and Mimura–Tsujikawa models–demonstrate that higher-order surrogate models significantly improve prediction accuracy while maintaining low computational cost. The proposed method provides a flexible, non-intrusive model reduction framework, well suited for the analysis of complex spatio-temporal pattern formation phenomena.

Data-driven modeling↗

Machine Learning-Based Process Control for Injection Molding of Recycled Polypropylene

The increased interest in artificial intelligence in manufacturing has driven the adoption of machine learning to optimize processes and improve efficiency. A key challenge in injection molding is the variability of recycled materials, which affects part quality and processing stability. This study presents a novel closed-loop process control approach for injection molding, leveraging machine learning to adaptively predict processing inputs and quality outcomes. The methodology was tested on five blends of recycled polypropylene (rPP), using artificial neural networks (ANNs), linear regression, and polynomial regression to model the relationships between material properties and process parameters. The dataset was split 80/20 into training and testing sets. The ANN model was implemented using TensorFlow and Keras, with six hidden layers of 32 neurons per layer, ReLU activation, and an Adam optimizer. Empirical tuning and early stopping were used to optimize performance and prevent overfitting. Predictions were evaluated based on mean absolute error (MAE), mean squared error (MSE), and percentage error. The results showed that yield stress, ultimate elongation, and part weight were accurately predicted within a 5% error for linear and polynomial regression models and within a 10% error for the ANN. However, modulus predictions were less reliable, with errors of ~11% for ANN and linear regression and ~40% for polynomial regression, reflecting the inherent variability of this property in rPP blends. Predictions of processing inputs had errors ranging from 3% to 25%, depending on the model and response variable. No single modeling approach was consistently superior across all responses, highlighting the complexity of the relationship between material properties, process parameters, and quality metrics. Overall, the work demonstrates that closed-loop process control, powered by machine learning, can effectively predict key quality parameters in injection molding of recycled materials. The proposed approach can improve process stability and material utilization, facilitating increased adoption of sustainable materials.

Krantz, Joshua↗

Automated and High-Throughput Phase Separation Control for Supramolecular Polymer Blends Enabled by Machine Learning

Supramolecular polymer blends (SPBs) offer tunable morphologies that dictate their macroscopic properties, yet their rational design is limited by the absence of predictive structure−morphology models. Here, we introduce a data-driven highthroughput workflow that integrates modular polymer synthesis, robotic formulation, automated morphology characterization, and machine learning (ML) for accelerated SPB discovery. Using a plug-and-play synthetic strategy, 33 hydrogen-bonding endfunctional homopolymers were prepared and orthogonally combined to generate 260 SPBs in 1 day. A fully automated atomic force microscopy (AFM) pipeline enabled systematic imaging, producing 2340 morphology data sets with minimal human intervention. Domain spacings were extracted through complementary imageprocessing methods and used to train ML models. A support vector regression (SVR) model accurately predicted target phase-separation sizes (50, 100, and 150 nm), which were experimentally validated. This work demonstrates the power of coupling high-throughput experimentation with ML to accelerate morphology discovery and provides one of the first large-scale experimental data sets for supramolecular polymer systems.

ML-guided polymer design↗

Ensemble Statistical Post-Processing of the National Air Quality Forecast Capability: Enhancing Ozone Forecasts in Baltimore, Maryland

An ensemble statistical post-processor (ESP) is developed for the National Air Quality Forecast Capability (NAQFC) to address the unique challenges of forecasting surface ozone in Baltimore, MD. Air quality and meteorological data were collected from the eight monitors that constitute the Baltimore forecast region. These data were used to build the ESP using a moving-block bootstrap, regression tree models, and extreme-value theory. The ESP was evaluated using a 10-fold cross-validation to avoid evaluation with the same data used in the development process. Results indicate that the ESP is conditionally biased, likely due to slight overfitting while training the regression tree models. When viewed from the perspective of a decision-maker, the ESP provides a wealth of additional information previously not available through the NAQFC alone. The user is provided the freedom to tailor the forecast to the decision at hand by using decision-specific probability thresholds that define a forecast for an ozone exceedance. Taking advantage of the ESP, the user not only receives an increase in value over the NAQFC, but also receives value for An ensemble statistical post-processor (ESP) is developed for the National Air Quality Forecast Capability (NAQFC) to address the unique challenges of forecasting surface ozone in Baltimore, MD. Air quality and meteorological data were collected from the eight monitors that constitute the Baltimore forecast region. These data were used to build the ESP using a moving-block bootstrap, regression tree models, and extreme-value theory. The ESP was evaluated using a 10-fold cross-validation to avoid evaluation with the same data used in the development process. Results indicate that the ESP is conditionally biased, likely due to slight overfitting while training the regression tree models. When viewed from the perspective of a decision-maker, the ESP provides a wealth of additional information previously not available through the NAQFC alone. The user is provided the freedom to tailor the forecast to the decision at hand by using decision-specific probability thresholds that define a forecast for an ozone exceedance. Taking advantage of the ESP, the user not only receives an increase in value over the NAQFC, but also receives value for

ozone↗

Version 8 SBUV Ozone Profile Trends Compared with Trends from a Zonally Averaged Chemical Model

Linear regression trends for the years 1979-2003 were computed using the new Version 8 merged Solar Backscatter Ultraviolet (SBUV) data set of ozone profiles. These trends were compared to trends computed using ozone profiles from the Goddard Space Flight Center (GSFC) zonally averaged coupled model. Observed and modeled annual trends between 50 N and 50 S were a maximum in the higher latitudes of the upper stratosphere, with southern hemisphere (SH) trends greater than northern hemisphere (NH) trends. The observed upper stratospheric maximum annual trend is -5.5 +/- 0.9 % per decade (1 sigma) at 47.5 S and -3.8 +/- 0.5 % per decade at 47.5 N, to be compared with the modeled trends of -4.5 +/- 0.3 % per decade in the SH and -4.0 +/- 0.2% per decade in the NH. Both observed and modeled trends are most negative in winter and least negative in summer, although the modeled seasonal difference is less than observed. Model trends are shown to be greatest in winter due to a repartitioning of chlorine species and the increasing abundance of chlorine with time. The model results show that trend differences can occur depending on whether ozone profiles are in mixing ratio or number density coordinates, and on whether they are recorded on pressure or altitude levels.

Rosenfield, Joan E.↗

The Impact of Atmospheric Dynamics and Anthropogenic Very Short-Lived Chlorine Species on the Recovery of Extra-Polar Ozone

The successful implementation of the Montreal Protocol has led to a decrease in the atmospheric abundance of ozone-depleting substances and a slowing in the destruction of the ozone layer. Previous studies have suggested that atmospheric dynamics, or a rise in compounds not regulated by the Montreal Protocol, such as very short-lived chlorine species (VSL Cl), may lead to a slower than expected recovery of the ozone layer. In this presentation, we examine the expected recovery of total column ozone (TCO) and stratospheric column ozone (SCO) to values observed in 1980 using a novel multiple linear regression (MLR) model that involves a month-by-month regression. The MLR model is trained to TCO anomalies from six data records (SBUV v8.7 MOD, SBUV v8.6 COH, WOUDC, GSG, GTO-ECV, MSR-2) over 1979 to 2021 for the Northern Hemisphere (35N – 60N), the Southern Hemisphere (60S – 35S) and the Tropics (20S – 20N), and SCO anomalies from ML-TOMCAT for these same zonal bands. The MLR includes the effect of halogens (equivalent effective stratospheric chlorine (EESC)), total solar irradiance, stratospheric aerosol optical depth, quasi-biennial oscillation, and the El Niño Southern Oscillation as regressors. We also include the effect of atmospheric dynamics such as the Brewer-Dobson Circulation, Arctic Oscillation, and the Antarctic Oscillation, and VSL Cl species in the formulation of EESC. We use a novel approach to the MLR framework, by separating the observed and regressor time series into the respective months (separating all of the Januarys, Februarys, etc.). Then we conduct a regression for each month from 1979 to 2021 for the three zonal bands denoted above. The model results for each month are combined together to achieve a full monthly time series from 1979 to 2021. We use this novel approach to ascertain the effect of atmospheric dynamics on TCO and SCO, since the dynamical proxies affect ozone in a distinctly different manner for various months. In this presentation, we will quantify the role of both the inclusion of VSL Cl species in the formulation of EESC as well as atmospheric dynamics in explaining the slower than expected recovery of extra-polar TCO and SCO over the past decade.

Laura A. McBride↗

Leveraging design of experiments to build chemometric models for the quantification of uranium (VI) and HNO3 by Raman spectroscopy

Partial least squares regression (PLSR) and support vector regression (SVR) models were optimized for the quantification of U(VI) (10–320 g L −1 ) and HNO 3 (0.6–6 M) by Raman spectroscopy with optimized calibration sets chosen by optimal design of experiments. The designed approach effectively minimized the number of samples in the calibration set for PLSR and SVR by selecting sample concentrations with a quadratic process model, despite complex confounding and covarying spectral features in the spectra. The top PLS2 model resulted in percent root mean square errors of prediction for U(VI), HNO 3 , and NO 3 − of 3.7%, 3.6%, and 2.9%, respectively. PLS1 models performed similarly despite modeling an analyte with a majority linear response (i.e., uranyl symmetric stretch) and another with more covarying vibrational modes (i.e., HNO 3 ). Partial least squares (PLS) model loadings and regression coefficients were evaluated to better understand the relationship between weaker Raman bands and covarying spectral features. Support vector machine models outperformed PLS1 models, resulting in percent root mean square error of prediction values for U(VI) and HNO 3 of 1.5% and 3.1%, respectively. The optimal nonlinear SVR model was trained using a similar number of samples (11) compared with the PLSR model, even though PLS is a linear modeling approach. The generic D-optimal design presented in this work provides a robust statistical framework for selecting training set samples in disparate two-factor systems. This approach reinforces Raman spectroscopy for the quantification of species relevant to the nuclear fuel cycle and provides a robust chemometric modeling approach to bolster online monitoring in challenging process environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Comparison of some biased estimation methods (including ordinary subset regression) in the linear model

Ridge, Marquardt's generalized inverse, shrunken, and principal components estimators are discussed in terms of the objectives of point estimation of parameters, estimation of the predictive regression function, and hypothesis testing. It is found that as the normal equations approach singularity, more consideration must be given to estimable functions of the parameters as opposed to estimation of the full parameter vector; that biased estimators all introduce constraints on the parameter space; that adoption of mean squared error as a criterion of goodness should be independent of the degree of singularity; and that ordinary least-squares subset regression is the best overall method.

Sidik, S. M.↗

Lateral-Directional Parameter Estimation on the X-48B Aircraft Using an Abstracted, Multi-Objective Effector Model

The problem of parameter estimation on hybrid-wing-body aircraft is complicated by the fact that many design candidates for such aircraft involve a large number of aerodynamic control effectors that act in coplanar motion. This adds to the complexity already present in the parameter estimation problem for any aircraft with a closed-loop control system. Decorrelation of flight and simulation data must be performed in order to ascertain individual surface derivatives with any sort of mathematical confidence. Non-standard control surface configurations, such as clamshell surfaces and drag-rudder modes, further complicate the modeling task. In this paper, time-decorrelation techniques are applied to a model structure selected through stepwise regression for simulated and flight-generated lateral-directional parameter estimation data. A virtual effector model that uses mathematical abstractions to describe the multi-axis effects of clamshell surfaces is developed and applied. Comparisons are made between time history reconstructions and observed data in order to assess the accuracy of the regression model. The Cram r-Rao lower bounds of the estimated parameters are used to assess the uncertainty of the regression model relative to alternative models. Stepwise regression was found to be a useful technique for lateral-directional model design for hybrid-wing-body aircraft, as suggested by available flight data. Based on the results of this study, linear regression parameter estimation methods using abstracted effectors are expected to perform well for hybrid-wing-body aircraft properly equipped for the task.

Ratnayake, Nalin A.↗

Assessment of a remote sensing-based model for predicting malaria transmission risk in villages of Chiapas, Mexico

A blind test of two remote sensing-based models for predicting adult populations of Anopheles albimanus in villages, an indicator of malaria transmission risk, was conducted in southern Chiapas, Mexico. One model was developed using a discriminant analysis approach, while the other was based on regression analysis. The models were developed in 1992 for an area around Tapachula, Chiapas, using Landsat Thematic Mapper (TM) satellite data and geographic information system functions. Using two remotely sensed landscape elements, the discriminant model was able to successfully distinguish between villages with high and low An. albimanus abundance with an overall accuracy of 90%. To test the predictive capability of the models, multitemporal TM data were used to generate a landscape map of the Huixtla area, northwest of Tapachula, where the models were used to predict risk for 40 villages. The resulting predictions were not disclosed until the end of the test. Independently, An. albimanus abundance data were collected in the 40 randomly selected villages for which the predictions had been made. These data were subsequently used to assess the models' accuracies. The discriminant model accurately predicted 79% of the high-abundance villages and 50% of the low-abundance villages, for an overall accuracy of 70%. The regression model correctly identified seven of the 10 villages with the highest mosquito abundance. This test demonstrated that remote sensing-based models generated for one area can be used successfully in another, comparable area.

Insect Vectors/growth & development↗

Parameter uncertainties for imperfect surrogate models in the low-noise regime

Abstract Bayesian regression determines model parameters by minimizing the expected loss, an upper bound to the true generalization error. However, this loss ignores model form error, or misspecification, meaning parameter uncertainties are significantly underestimated and vanish in the large data limit. As misspecification is the main source of uncertainty for surrogate models of low-noise calculations, such as those arising in atomistic simulation, predictive uncertainties are systematically underestimated. We analyze the true generalization error of misspecified, near-deterministic surrogate models, a regime of broad relevance in science and engineering. We show that posterior parameter distributions must cover every training point to avoid a divergence in the generalization error and design a compatible ansatz which incurs minimal overhead for linear models. The approach is demonstrated on model problems before application to thousand-dimensional datasets in atomistic machine learning. Our efficient misspecification-aware scheme gives accurate prediction and bounding of test errors in terms of parameter uncertainties, allowing this important source of uncertainty to be incorporated in multi-scale computational workflows.

Swinburne, Thomas D. (ORCID:0000000232554257)↗

Time-dependent oral absorption models

The plasma concentration-time profiles following oral administration of drugs are often irregular and cannot be interpreted easily with conventional models based on first- or zero-order absorption kinetics and lag time. Six new models were developed using a time-dependent absorption rate coefficient, ka(t), wherein the time dependency was varied to account for the dynamic processes such as changes in fluid absorption or secretion, in absorption surface area, and in motility with time, in the gastrointestinal tract. In the present study, the plasma concentration profiles of propranolol obtained in human subjects following oral dosing were analyzed using the newly derived models based on mass balance and compared with the conventional models. Nonlinear regression analysis indicated that the conventional compartment model including lag time (CLAG model) could not predict the rapid initial increase in plasma concentration after dosing and the predicted Cmax values were much lower than that observed. On the other hand, all models with the time-dependent absorption rate coefficient, ka(t), were superior to the CLAG model in predicting plasma concentration profiles. Based on Akaike's Information Criterion (AIC), the fluid absorption model without lag time (FA model) exhibited the best overall fit to the data. The two-phase model including lag time, TPLAG model was also found to be a good model judging from the values of sum of squares. This model also described the irregular profiles of plasma concentration with time and frequently predicted Cmax values satisfactorily. A comparison of the absorption rate profiles also suggested that the TPLAG model is better at prediction of irregular absorption kinetics than the FA model. In conclusion, the incorporation of a time-dependent absorption rate coefficient ka(t) allows the prediction of nonlinear absorption characteristics in a more reliable manner.

Non-NASA Center↗