Search NASA⌕ Search

SEARCH · Search NASA

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Spatially Refined Satellite Gravimetry Captures Human Signatures in Global Terrestrial Water Storage Trends

Human activities have directly altered the water cycle through water management, aquifer pumping, agricultural irrigation, and land use change. Although satellite gravimetry has transformed global hydrological research, its coarse resolution limits attribution of freshwater change to human activities at many management-relevant scales. Here we assessed global terrestrial water storage (TWS) trends from April 2002 to November 2025 using “stacked” regression of Level-1B intersatellite ranging data, which leverages temporal information and variability to dramatically improve effective spatial resolution relative to standard approaches. We combined this refined product with rigorous uncertainty analysis, autocorrelation-robust geostatistical methods, and literature assessment to evaluate TWS trend associations with land and water use, climate variability, and glacial mass loss. We identified TWS trend hotspots exhibiting significant spatial associations with anthropogenic 40 drivers, including groundwater and surface-water irrigation, rainfed agriculture, deforestation, and reservoir impoundment. Across these regions, cumulative TWS losses (3,122 Gt) substantially exceeded gains (2,432 Gt). Compared with traditional regression of monthly mascons, our approach yielded regional trend magnitudes that are on average 33% larger, revealing that global freshwater depletion, particularly from groundwater pumping, is considerably more acute than previously estimated. Multivariate regression models show that humans account for a significant share of the spatial variability in TWS trends on every non-polar continent except Australia. We detected localized TWS gains linked to rainfed agriculture, surface water irrigation, and reservoir filling that were unresolved in earlier gravimetric studies. The methodology provides a foundation for future gravity missions to independently track decadal freshwater change with unprecedented spatial fidelity.

groundwater↗

Quantile regression-enriched event modeling framework for dropout analysis in high-temperature superconductor manufacturing

High-temperature superconductor (HTS) tapes have shown promising characteristics of high critical current, which are prerequisites for applications in high-field magnets. Due to the unstable growth conditions in the HTS manufacturing process, however, the frequent occurrences of dropouts in the critical current impede the consistent performance of HTS tapes. To manufacture HTS tapes with large scale, high yield, and uniform performance, it is essential to develop novel data analysis approaches for modeling the dropouts and identifying the related important process parameters. Conventional methods for modeling recurrent events, such as the point process, require the extraction of events from quality measurements. As the critical current is a continuous process, it may not comprehensively represent the drop patterns by transforming the time-series measurements into a set of events. Here, to solve this issue, we develop a novel quantile regression-enriched event modeling (QREM) framework that integrates the non-homogeneous Poisson process for modeling the occurrence of dropouts and the quantile regression for capturing the drop patterns. By incorporating the feature selection and regularization, the proposed framework identifies a set of significant process parameters that can potentially cause the dropouts of HTS tapes. The proposed method is tested on real HTS tapes produced using an advanced manufacturing process, successfully identifying important parameters that influence dropout events including the substrate temperature and voltage. The results demonstrate that the proposed QREM method outperforms the standard point process in predicting the occurrence of dropouts.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A Method for Calculating the Probability of Successfully Completing a Rocket Propulsion Ground Test

Propulsion ground test facilities face the daily challenge of scheduling multiple customers into limited facility space and successfully completing their propulsion test projects. Over the last decade NASA s propulsion test facilities have performed hundreds of tests, collected thousands of seconds of test data, and exceeded the capabilities of numerous test facility and test article components. A logistic regression mathematical modeling technique has been developed to predict the probability of successfully completing a rocket propulsion test. A logistic regression model is a mathematical modeling approach that can be used to describe the relationship of several independent predictor variables X(sub 1), X(sub 2),.., X(sub k) to a binary or dichotomous dependent variable Y, where Y can only be one of two possible outcomes, in this case Success or Failure of accomplishing a full duration test. The use of logistic regression modeling is not new; however, modeling propulsion ground test facilities using logistic regression is both a new and unique application of the statistical technique. Results from this type of model provide project managers with insight and confidence into the effectiveness of rocket propulsion ground testing.

Messer, Bradley↗

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE↗

Mechanistic Modeling of TEG Dehydrator Emissions in Oil and Gas Industry

This work presents a mechanistic modeling approach for simulating methane emissions from triethylene glycol (TEG) dehydrators used in oil & gas (O&G) operations. The model was developed as a modular component of the Mechanistic Air Emissions Simulator (MAES) tool, incorporating species-specific absorption and emission dynamics through two-level, second-order polynomial regression (PR) models trained on ProMax simulation data: (1) species-level regression models that track the transfer rates of individual gas species within the dehydrator unit streams, and (2) outlet flow stream regression models that predict the fraction of inlet gas distributed among the outlet streams of the dehydrator unit. These behaviors were characterized over a range of glycol circulation ratios, wet gas pressures, and temperatures. The model was validated using root mean square error (RMSE) analysis. The species-level PR achieved low root mean square error (RMSE) values (<0.03) for light hydrocarbon species across all dehydrator components, ranging from 0.0009 for methane to 0.029 for normal pentane. Similarly, the outlet-level PR yielded RMSE values below 0.002 for the dry gas fraction, 0.001 for the flash tank fraction, and 0.002 for the still vent fraction, demonstrating strong agreement between predicted and reference ProMax values. When deployed at field facilities, the model significantly improved MAES-simulated dehydrator emissions, revealing that gas-assisted glycol pump emissions are the dominant contributors to both dehydrator-level and site-level methane emissions under uncontrolled conditions. Further analysis of the 154 dehydrator units reported by operators under the AMI 2024 project showed that 54 units (31%) used gas-driven glycol pumps, of which 6 units (11%) operated with uncontrolled flash tanks, and 22 units (40.7%) were identified as potentially oversized. Of the six dehydrator units with uncontrolled gas-assisted pumps, pump emissions accounted for 90.25% of total dehydrator emissions and 63.10% of total site-level emissions. These findings highlight substantial opportunities for emissions mitigation through equipment upgrades.

MAES↗

Predicting Ground Delay Program at an Airport Based on Meteorological Conditions

In this paper, we present two supervised-learning models, logistic regression and decision tree, to predict occurrence of ground delay program at an airport based on meteorological conditions and scheduled traffic demand. Such predictive capabilities can help the Federal Aviation Administration traffic managers and airline dispatchers to prepare mitigation strategies to reduce the impact of adverse weather. The models are applied to predict ground delay program occurrence at two major U.S. airports: Newark Liberty Intl. and San Francisco Intl. airports. The logistic regression model estimates the probability that a ground delay program will occur during a given hour. Decision tree, on the other hand, classifies an hour as a ground delay program or not based on the input variables. Results indicate that both models perform significantly better than a purely random prediction of ground delay program occurrence at the two airports. The logistic regression model performs better than the decision tree model. The degree to which various input variables impact the probability of ground delay program vary between the two airports. While the enroute convective weather is a dominant factor causing ground delay programs at New York airports, poor visibility and low cloud ceiling caused by marine stratus are major drivers of ground delay programs at San Francisco Intl. airport.

traffic flow management↗

Multiobjective Constrained Symbolic Regression for Predictive Modeling of Material Creep Behavior

When creep testing is repeated on samples of the same alloy under the same parametric conditions (i.e., stress and temperature), the resulting strain/time curves can vary from each other considerably as shown in Figure 1 [1]. The time required to creep test a material to rupture can extend to the order of years. Because of this, a numerical model that can quickly analyze the incomplete results of an ongoing experiment to predict 1) the incomplete portion of the strain/time curve leading up to the rupture point and 2) the rupture point itself would be of great utility to the materials community. Such a model has the potential to save 1) the time required to finish running the experiment to rupture 2) the associated monetary cost of finishing said experiment. Furthermore, it would be advantageous if the predictive model could give a parametric function modeling strain/time curves for material scientists to investigate the impact of the temperature and stress parameters on the resulting creep behavior. This work introduces a piecewise symbolic regression algorithm to predict the remainder of the strain/time curve. Preliminary results show good model performance.

36 MATERIALS SCIENCE↗

Thermal Testing and Analysis Techniques for Wires and Wire Bundles

Determining wire and wire bundle amperage capacity (i.e., “ampacity”) currently relies on the use of standards to derate wire ampacity when in a bundle configuration. The feasibility of developing physics-based and regression thermal models of single wires and wire bundles to determine ampacity using a customized test apparatus was investigated during a pathfinder study. A test facility was developed and various wire and wire bundle articles were tested under a variety of temperature and pressure conditions using an efficient test matrix formulated using Design of Experiments (DOE) techniques. Physics-based models were developed and correlated to the test results. Regression models were formulated and compared to test results and standards.

Thermal↗

Recursive Algorithm For Linear Regression

Order of model determined easily. Linear-regression algorithhm includes recursive equations for coefficients of model of increased order. Algorithm eliminates duplicative calculations, facilitates search for minimum order of linear-regression model fitting set of data satisfactory.

Varanasi, S. V.↗

Decayheatml

This code is designed to predict and analyze the decay heat generated in molten salt reactors (MSRs) using a hybrid approach that combines machine learning and segmented polynomial fitting. The accurate prediction of decay heat is essential for reactor safety and the optimization of spent fuel storage. The code operates through several key components: 1) Data Architecture: It incorporates a modular data architecture that handles various MSR-specific operational parameters such as power density, humidity content, and air ingress. These parameters are sampled using Sobol sequences to ensure comprehensive coverage of operational uncertainties. 2) Machine Learning Framework: The code employs a diverse set of machine learning models, including polynomial regression, decision trees, random forests, gradient boosting, support vector regression, k-nearest neighbors, multi-layer perceptrons, and symbolic regression. These models are trained to predict decay heat over a wide temporal range, from immediate shutdown up to 10,000 years. 3) Region-Optimized Training: The temporal domain is divided into multiple regions, each modeled separately to capture distinct decay heat characteristics across different time scales. This approach significantly improves the accuracy and interpretability of predictions. 4) Segmented Polynomial Interpretation (SPI): The SPI method translates machine learning predictions into piecewise polynomial equations. These equations are physically interpretable and can be directly integrated into existing engineering workflows and safety analyses. 5) Front-End Interfaces: The code includes both a Jupyter notebook interface for research development and a Streamlit web application for operational deployment. These interfaces allow users to interactively explore decay heat predictions, adjust operational parameters, and visualize results in real-time. 6) Applications: The framework supports various applications, including safety system validation and spent fuel container optimization. It enables real-time evaluation of worst-case decay heat scenarios, informing the design of passive safety systems and optimizing container designs for long-term storage. Overall, this code provides a robust, accurate, and user-friendly tool for predicting decay heat in MSRs, enhancing reactor safety, and optimizing spent fuel management.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

Modeling the Height of Young Forests Regenerating from Recent Disturbances in Mississippi using Landsat and ICESat data

Many forestry and earth science applications require spatially detailed forest height data sets. Among the various remote sensing technologies, lidar offers the most potential for obtaining reliable height measurement. However, existing and planned spaceborne lidar systems do not have the capability to produce spatially contiguous, fine resolution forest height maps over large areas. This paper describes a Landsat-lidar fusion approach for modeling the height of young forests by integrating historical Landsat observations with lidar data acquired by the Geoscience Laser Altimeter System (GLAS) instrument onboard the Ice, Cloud, and land Elevation (ICESat) satellite. In this approach, "young" forests refer to forests reestablished following recent disturbances mapped using Landsat time-series stacks (LTSS) and a vegetation change tracker (VCT) algorithm. The GLAS lidar data is used to retrieve forest height at sample locations represented by the footprints of the lidar data. These samples are used to establish relationships between lidar-based forest height measurements and LTSS-VCT disturbance products. The height of "young" forest is then mapped based on the derived relationships and the LTSS-VCT disturbance products. This approach was developed and tested over the state of Mississippi. Of the various models evaluated, a regression tree model predicting forest height from age since disturbance and three cumulative indices produced by the LTSS-VCT method yielded the lowest cross validation error. The R(exp 2) and root mean square difference (RMSD) between predicted and GLAS-based height measurements were 0.91 and 1.97 m, respectively. Predictions of this model had much higher errors than indicated by cross validation analysis when evaluated using field plot data collected through the Forest Inventory and Analysis Program of USDA Forest Service. Much of these errors were due to a lack of separation between stand clearing and non-stand clearing disturbances in current LTSS-VCT products and difficulty in deriving reliable forest height measurements using GLAS samples when terrain relief was present within their footprints. In addition, a systematic underestimation of about 5 m by the developed model was also observed, half of which could be explained by forest growth that occurred between field measurement year and model target year. The remaining difference suggests that tree height measurements derived using waveform lidar data could be significantly underestimated, especially for young pine forests. Options for improving the height modeling approach developed in this study were discussed.

Li, Ainong↗

A Data-Driven Method for Modeling Creep-Fatigue Stress- Strain Behavior Using Neural ODEs

In this paper, we introduce a data-driven machine learning approach for modeling one-dimensional stress–strain behavior under cyclic loading, utilizing experimental data from the nickel-based Alloy 617. The study employs uniaxial creep–fatigue test data acquired under various loading histories and compares two distinct neural network-based ODE models. The first model, known as the black-box model, comprehensively describes the strain–stress relationship using a Neural ODE equation. To interpret this black-box model, we apply the Sparse Identification of Nonlinear Dynamical Systems (SINDy) technique, transforming the black-box model into an equation-based model using symbolic regression. The second model, the Neural flow rule model, incorporates Hooke’s Law for the linear elastic component, with the nonlinear part characterized by a Neural ODE. Both models are trained with experimental data to accurately reflect the observed stress–strain behavior. We conduct a detailed comparison with the standard Chaboche model, which includes three back stresses. Our results demonstrate that the neural network-based ODE models precisely capture the experimental creep–fatigue mechanical behavior, exceeding the standard Chaboche model’s accuracy. Furthermore, an interpretable model derived from the black-box neural ODE model through symbolic regression achieves accuracy comparable to the Chaboche model, enhancing its interpretability. The results highlight the potential of neural network-based ODE models to depict complex creep–fatigue behavior, eliminating the necessity for experts to define a specific, material-focused model form.

creep-fatigue↗

Efficient data-driven regression for reduced-order modeling of spatial pattern formation

We present an efficient data-driven regression approach for constructing reduced-order models (ROMs) of reaction-diffusion systems exhibiting pattern formation. The ROMs are learned non-intrusively from available training data of physically accurate numerical simulations. The method can be applied to general nonlinear systems through the use of polynomial model form, while not requiring knowledge of the underlying physical model, governing equations, or numerical solvers. The process of learning ROMs is posed as a low-cost least-squares problem in a reduced-order subspace identified via Proper Orthogonal Decomposition (POD). Numerical experiments on classical pattern-forming systems–including the Schnakenberg and Mimura–Tsujikawa models–demonstrate that higher-order surrogate models significantly improve prediction accuracy while maintaining low computational cost. The proposed method provides a flexible, non-intrusive model reduction framework, well suited for the analysis of complex spatio-temporal pattern formation phenomena.

Data-driven modeling↗

Machine Learning-Based Process Control for Injection Molding of Recycled Polypropylene

The increased interest in artificial intelligence in manufacturing has driven the adoption of machine learning to optimize processes and improve efficiency. A key challenge in injection molding is the variability of recycled materials, which affects part quality and processing stability. This study presents a novel closed-loop process control approach for injection molding, leveraging machine learning to adaptively predict processing inputs and quality outcomes. The methodology was tested on five blends of recycled polypropylene (rPP), using artificial neural networks (ANNs), linear regression, and polynomial regression to model the relationships between material properties and process parameters. The dataset was split 80/20 into training and testing sets. The ANN model was implemented using TensorFlow and Keras, with six hidden layers of 32 neurons per layer, ReLU activation, and an Adam optimizer. Empirical tuning and early stopping were used to optimize performance and prevent overfitting. Predictions were evaluated based on mean absolute error (MAE), mean squared error (MSE), and percentage error. The results showed that yield stress, ultimate elongation, and part weight were accurately predicted within a 5% error for linear and polynomial regression models and within a 10% error for the ANN. However, modulus predictions were less reliable, with errors of ~11% for ANN and linear regression and ~40% for polynomial regression, reflecting the inherent variability of this property in rPP blends. Predictions of processing inputs had errors ranging from 3% to 25%, depending on the model and response variable. No single modeling approach was consistently superior across all responses, highlighting the complexity of the relationship between material properties, process parameters, and quality metrics. Overall, the work demonstrates that closed-loop process control, powered by machine learning, can effectively predict key quality parameters in injection molding of recycled materials. The proposed approach can improve process stability and material utilization, facilitating increased adoption of sustainable materials.

Krantz, Joshua↗

Automated and High-Throughput Phase Separation Control for Supramolecular Polymer Blends Enabled by Machine Learning

Supramolecular polymer blends (SPBs) offer tunable morphologies that dictate their macroscopic properties, yet their rational design is limited by the absence of predictive structure−morphology models. Here, we introduce a data-driven highthroughput workflow that integrates modular polymer synthesis, robotic formulation, automated morphology characterization, and machine learning (ML) for accelerated SPB discovery. Using a plug-and-play synthetic strategy, 33 hydrogen-bonding endfunctional homopolymers were prepared and orthogonally combined to generate 260 SPBs in 1 day. A fully automated atomic force microscopy (AFM) pipeline enabled systematic imaging, producing 2340 morphology data sets with minimal human intervention. Domain spacings were extracted through complementary imageprocessing methods and used to train ML models. A support vector regression (SVR) model accurately predicted target phase-separation sizes (50, 100, and 150 nm), which were experimentally validated. This work demonstrates the power of coupling high-throughput experimentation with ML to accelerate morphology discovery and provides one of the first large-scale experimental data sets for supramolecular polymer systems.

ML-guided polymer design↗