Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Determination of confinement regime boundaries via separatrix parameters on Alcator C-Mod based on a model for interchange-drift-Alfvén turbulence

The separatrix operational space (SepOS) model (Eich and Manz 2021 Nucl. Fusion 61 086017) is shown to predict the L–H transition, the L-mode density limit, and the ideal magnetohydrodynamic ballooning limit in terms of separatrix parameters for a wide range of Alcator C-Mod plasmas. The model is tested using Thomson scattering measurements across a wide range of operating conditions on C-Mod, spanning $\overline{n}_{\mathrm{e}}$ = 0.3–5.5 $\times 10^{20}\,$m −3 , $B_{\mathrm{t}} = 2.5$–8.0 T, and $B_{\mathrm{p}}$= 0.1–1.2 T. An empirical regression for the electron pressure gradient scale length, $\lambda_{{p}_{\mathrm{e}}}$, against a turbulence control parameter, $\alpha_{\mathrm{t}}$, and the poloidal fluid gyroradius, $\rho_{\mathrm{s,p}}$, is constructed for H-modes and found to require positive exponents for both regression parameters, indicating turbulence widening of near-scrape-off layer widths at high $\alpha_{\mathrm{t}}$ and an inverse scaling with $B_{\mathrm{p}}$, consistent with results on ASDEX Upgrade. The SepOS model is also tested in the unfavorable drift direction and found to apply well to all three boundaries, including the L–H transition as long as a correction to the Reynolds energy transfer term, $\alpha_\mathrm{RS} \lt 1$ is applied. I-modes typically exist in the unfavorable drift direction for values of $\alpha_{\mathrm{t}} \lesssim 0.3$. Finally, an experiment studying the transition between the Type-I ELMy and EDA H-mode is analyzed using the same framework. It is found that a recently identified boundary $\alpha_{\mathrm{t}} = 0.55$ at excludes most EDA H-modes but that the balance of wavenumbers responsible for the L-mode density limit, namely $k_\mathrm{EM} = k_\mathrm{RBM}$, may better describe the transition on C-Mod. The ensemble of boundaries validated and explored is then applied to project regime access and limit avoidance for the SPARC primary reference discharge parameters.

ELM suppression↗

Sparse Cholesky factorization for solving nonlinear PDEs via Gaussian processes

In recent years, there has been widespread adoption of machine learning-based approaches to automate the solving of partial differential equations (PDEs). Among these approaches, Gaussian processes (GPs) and kernel methods have garnered considerable interest due to their flexibility, robust theoretical guarantees, and close ties to traditional methods. They can transform the solving of general nonlinear PDEs into solving quadratic optimization problems with nonlinear, PDE-induced constraints. However, the complexity bottleneck lies in computing with dense kernel matrices obtained from pointwise evaluations of the covariance kernel, and its partial derivatives, a result of the PDE constraint and for which fast algorithms are scarce. The primary goal of this paper is to provide a near-linear complexity algorithm for working with such kernel matrices. We present a sparse Cholesky factorization algorithm for these matrices based on the near-sparsity of the Cholesky factor under a novel ordering of pointwise and derivative measurements. The near-sparsity is rigorously justified by directly connecting the factor to GP regression and exponential decay of basis functions in numerical homogenization. We then employ the Vecchia approximation of GPs, which is optimal in the Kullback-Leibler divergence, to compute the approximate factor. This enables us to compute ϵ-approximate inverse Cholesky factors of the kernel matrices with complexity O(N log d (N/ϵ)) in space and O(N log 2d (N/ϵ)) in time. We integrate sparse Cholesky factorizations into optimization algorithms to obtain fast solvers of the nonlinear PDE. We numerically illustrate our algorithm’s near-linear space/time complexity for a broad class of nonlinear PDEs such as the nonlinear elliptic, Burgers, and Monge-Ampère equations. In summary, we provide a fast, scalable, and accurate method for solving general PDEs with GPs and kernel methods.

97 MATHEMATICS AND COMPUTING↗

CovTransformer: A transformer model for SARS-CoV-2 lineage frequency forecasting

With hundreds of SARS-CoV-2 lineages circulating in the global population, there is an ongoing need for predicting and forecasting lineage frequencies and thus identifying rapidly expanding lineages. Accurate prediction would allow for more focused experimental efforts to understand pathogenicity of future dominating lineages and characterize the extent of their immune escape. Here, we first show that the inherent noise and biases in lineage frequency data make a commonly-used regression-based approach unreliable. To address this weakness, we constructed a machine learning model for SARS-CoV-2 lineage frequency forecasting, called CovTransformer, based on the transformer architecture. We designed our model to navigate challenges such as a limited amount of data with high levels of noise and bias. We first trained and tested the model using data from the UK and the USA, and then tested the generalization ability of the model to many other countries and US states. Remarkably, the trained model makes accurate predictions two months into the future with high levels of accuracy both globally (in 31 countries with high levels of sequencing effort) and at the US-state level. Our model performed substantially better than a widely used forecasting tool, the multinomial regression model implemented in Nextstrain, demonstrating its utility in SARS-CoV-2 monitoring. Assuming a newly emerged lineage is identified and assigned, our test using retrospective data shows that our model is able to identify the dominating lineages 7 weeks in advance on average before they became dominant. Overall, our work demonstrates that transformer models represent a promising approach for SARS-CoV-2 forecasting and pandemic monitoring.

60 APPLIED LIFE SCIENCES↗

Estradiol associations with brain functional connectivity in postmenopausal women

Abstract Objective Previous studies have found that estrogens play a role in functional connectivity in the brain; however, little research has been done regarding how estradiol is associated with functional connectivity in postmenopausal women. The purpose of this study was to examine the relationship between estradiol and functional connectivity in postmenopausal women. Methods Structural and blood oxygenation level–dependent resting-state magnetic resonance imaging scans of 88 cognitively healthy postmenopausal individuals were obtained along with blood samples collected the same day as the magnetic resonance imaging to assess hormone levels. We generated connectivity values in CONN toolbox version 20.b, an SPM-based software. Results A regression analysis was run using estradiol level and regions of interest (ROI), including the hippocampus, parahippocampus, dorsolateral prefrontal cortex, and precuneus. Estradiol level was found to enhance parahippocampal gyrus anterior division left functional connectivity during ROI-to-ROI regression analysis. Estradiol enhanced functional connectivity between the parahippocampal gyrus anterior division left and the precuneus as well as the parahippocampal gyrus anterior division left and parahippocampal gyrus posterior division right. An exploratory analysis showed that years since the final menstrual period was related to enhanced connectivity between regions within the frontoparietal network. Conclusions These results illustrated the relationship between estradiol level and functional connectivity in postmenopausal women. They have implications for understanding how the functioning of the brain changes for individuals after menopause that may eventually lead to changes in cognition and behavior in older ages.

Obstetrics & Gynecology↗

Gaussian-process generative model for the QCD equation of state

We develop a generative model for the nuclear matter equation of state at zero net baryon density using the Gaussian process regression method. We impose first-principles theoretical constraints from lattice quantum chromodynamics and hadron resonance gas at high- and low-temperature regions, respectively. By allowing the trained Gaussian process regression model to vary freely near the phase transition region, we generate random smooth crossover equations of state with different speeds of sound that do not rely on specific parametrizations. Here, we explore a collection of experimental observable dependencies on the generated equations of state, which paves the groundwork for future Bayesian inference studies to use experimental measurements from relativistic heavy-ion collisions to constrain the nuclear matter equation of state.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Flux Hypothesis for Odd Transport Phenomena

Onsager's regression hypothesis makes a fundamental connection between macroscopic transport phenomena and the average relaxation of spontaneous microscopic fluctuations. This relaxation, however, is agnostic to odd transport phenomena, in which fluxes run orthogonal to the gradients driving them. To account for odd transport, we generalize the regression hypothesis, postulating that macroscopic linear constitutive laws are, on average, obeyed by microscopic fluctuations, whether they contribute to relaxation or not. From this "flux hypothesis," Green-Kubo and reciprocal relations follow, elucidating the separate roles of broken time-reversal and parity symmetries underlying various odd transport coefficients. As an application, we derive and verify the Green-Kubo relation for odd collective diffusion in chiral active matter, first in an analytically tractable model and subsequently through molecular dynamics simulations of concentrated active spinners.

Hargus, Cory↗

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

What Shapes Transportation Charging Infrastructure Availability? Evidence from Tennessee

This study examines how community, travel, and freight characteristics relate to public charging infrastructure availability across Tennessee ZIP codes. We link Alternative Fuels Data Center station locations with traffic, socioeconomic, demographic, commuting, and freight employment data to build a ZIP code-level dataset. Ordinary least squares regression captures variation in chargers per 10,000 residents (R2=0.311). Quantile regressions at the 25th, 50th, and 75th percentiles, with pseudo R2 values up to 0.099, show that the determinants of infrastructure availability differ across low-, medium-, and high-availability areas. Percent female, percent car commuters, average household size, and median age are negatively associated with charging availability across much of the distribution. Truck traffic is positively associated only in lower-availability ZIP codes, while vehicle miles traveled shifts from a negative association at the lower end of the distribution to a positive association at the upper end. The results provide insight into how public charging deployment aligns with community characteristics, mobility demand, and freight activity across Tennessee. Future work can distinguish charger types and power levels, incorporate land-use and temporal rollout patterns, and examine how charging infrastructure needs differ across urban and rural contexts.

Calderón, Oriana [University of Tennessee, Knoxvil↗

Data-Driven Digital Twin for Reliability Assessment of DC/DC Buck Converter

In commercial applications, the operation of DC/DC converters significantly impacts overall system performance and long-term reliability. This study introduces a data-driven digital twin (DT) approach for estimating critical degradation parameters of DC/DC BUCK converter under steady-state condition. Initially, a circuit-level MATLAB/Simulink digital model (DM C ) is refined against a hardware prototype’s switching model dataset using offline particle swarm optimization. The optimized digital model’s steady-state response is then verified with its average model response while varying the duty and load. Subsequently, degradation profiles are imposed on the inductor, capacitor, MOSFET in the DMC. A large dataset is generated from this model, allowing training, validation, and testing of machine learning (ML) models for component health regression tasks. The proposed method employs random forest ML models, achieving impressive regression results with a squared R value as high as 0.99978 and a root mean square error of 4.2× 10 –6 . The method is further validated on a medium power level DC/DC BUCK prototype with varying load conditions, and includes the analysis of MOSFET’s on-resistance under degradation conditions. This data-driven DT method shows promise for identifying parasitic degradation and ohmic loss parameters, enhancing converter reliability assessments in a non-invasive, generalized, and computationally efficient manner.

14 SOLAR ENERGY↗

Overview of RFID Applications Utilizing Neural Networks

As Radio Frequency Identification (RFID) methods continue to evolve to higher levels of complexity, one form of machine learning is making its appearance. The use of Neural Networks (NN) in the RFID field is steadily increasing, and in the fields of localization and activity recognition, promising results are being shown from a variety of research. RFID applications fall primarily under two types of problems including regression and classification. We analyze RIFD localization techniques which fall under regression, and activity recognition which falls under classification. Many works don’t classify themselves as activity recognition methods, but because they fall under the classification category, we still consider them as activity recognition techniques. This research overviews the Neural Network models in the localization field based on whether they can perform independently of the environment in which they were tested. For activity recognition and accessory fields, the major methods involve tag-based and tag-free approaches. In conclusion, after the models are surveyed, a comparison study is given to examine what may be the cause for increased accuracy between different Neural Network models.

42 ENGINEERING↗

Robust Design Under Uncertainty in Quantum Error Mitigation

Error mitigation techniques are crucial to achieving near-term quantum advantage. Classical postprocessing of quantum computation outcomes is a popular approach for error mitigation, which includes methods, such as zero noise extrapolation, virtual distillation, and learning-based error mitigation. However, these techniques have limitations due to the propagation of uncertainty resulting from the finite shot number of a quantum measurement. In this work, we introduce general and unbiased methods for quantifying the uncertainty and error of error-mitigated observables based on the strategic sampling of error mitigation outcomes. We then extend our approach to demonstrate the optimization of performance and robustness of error mitigation under uncertainty. To illustrate our methods, we apply them to zero noise extrapolation and Clifford date regression in the ground state of the XY model simulated using depolarizing and International Business Machines Corporation (IBM) Toronto noise models, respectively. In particular, we optimize the choice of noise levels and the allocation of shots for zero noise extrapolation and the distribution of the training circuits for Clifford data regression. While our methods are readily applicable to any postprocessing-based error mitigation approach, in practice they must not be prohibitively expensive—even though they perform optimizations of the error mitigation hyperparameters requiring sampling of a statistical distribution of error mitigation outcomes. By leveraging surrogate-based optimization, we show that our methods can efficiently perform optimal design for a zero noise extrapolation implementation. We then further demonstrate the transferability of learned zero noise extrapolation hyperparameters to other similar circuits.

97 MATHEMATICS AND COMPUTING↗

DEPRECATED AI-Batt-OS (Autonomous Identification of Battery Life Models - Open Source) [SWR 21-17]

DEPRECATED. This repository was archived by the owner on Jun 30, 2026. It is now read-only. Open source implementation of some of the methods utilized by AI-Batt, a battery lifetime modeling and analysis toolkit provided by the National Laboratory of the Rockies (NLR). This software demonstrates the use of bi-level optimization and symbolic regression techniques to semi-autonomously identify algebraic models predicting the capacity fade of lithium-ion batteries during calendar aging. Modeling the degradation of batteries is a complex task, due to the difficulty in separating the time-dependent and time-independent factors impacting cell level degradation, across multiple data series with different numbers of measurements and/or data quality. Bi-level optimization enables model parameters to be optimized to either the entire data set or to individual data series, allowing statistical disambiguation of global behaviors (data series independent) and local behaviors (data series dependent). Symbolic regression is used to automatically search for optimal low-dimesional models predicting the variation of locally optimized parameters versus time-independent experimental variables from millions of possible models, resulting in a more accurate and repeatable model identification process than is possible by a manual search. The provided tools also implement cross-validation and bootstrap resampling schemes, empowering statistical model comparison/selection and quantification of model uncertainties. An example script replicates the results from the manuscript "Challenging Practices of Algebraic Battery Life Models through Statistical Validation and Model Identification via Machine-Learning", submitted to ECS. All code is written in MATLAB. Requires the Statistics and Machine Learning Toolbox. Contact Dr. Paul Gasper at Paul.Gasper@nlr.gov for any questions.

Gasper, Paul [National Renewable Energy Lab. (NREL↗

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR↗

Moltensaltpropnet

MoltenSaltPropnet is a physics-informed machine learning framework that aims to predict the thermophysical properties of molten fluoride and chloride salt mixtures, which are crucial for the design and safety of Generation IV molten salt reactors. The code processes data from the Molten-Salt Thermal Properties Database (MSTDB-TP) and the Janz compendium, converting critically evaluated correlations into fast, differentiable surrogate models for density, viscosity, thermal conductivity, and heat capacity across 448 distinct salt systems. The implementation consists of several key components: 1. Data Curation: The code parses and cleans the raw data, normalizing elemental mole fractions and extracting relevant regression coefficients for various thermophysical properties. 2. Feature Engineering: It generates fixed-length numerical descriptors that encapsulate the composition and temperature, incorporating polynomial interaction terms and dimensionality-reduction techniques to optimize model performance. 3. Coefficient Learning: Four different machine learning architectures are employed: a deep residual network (ResNet), a Kolmogorov–Arnold network (KAN), a sparsity-inducing neural network (SNN), and classical regression models. Each model learns to predict coefficients that define the temperature-dependent correlations for the thermophysical properties. 4. Property Reconstruction: The predicted coefficients are used to compute temperature-dependent property values, ensuring positivity and monotonic trends through a composite loss function that enforces physical constraints. 5. User Interface: An open-source web application enables users to filter the database, train task-specific models, and visualize the results, allowing for rapid exploration of candidate salt mixtures. MoltenSaltPropnet bridges the gap between limited experimental data and high-fidelity reactor simulations, providing a powerful tool for researchers in the field of molten salt reactors and advanced nuclear energy systems.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

RxnRover/amlro

AMLRO (Active Machine Learning Reaction Optimizer) is an open-source framework designed to accelerate chemical reaction optimization using active learning with classical machine learning regression models. AMLRO integrates space-filling sampling strategies (e.g., Sobol and Latin Hypercube sampling) with iterative model training, prediction, and experiment selection to efficiently navigate complex reaction spaces. The platform supports multiple regression models, flexible multi-objective definitions, and user-defined parameter bounds, enabling data-efficient optimization from small initial datasets. AMLRO is designed for ease of use by experimentalists and can operate as a standalone decision-support tool or be integrated into closed-loop automated experimentation workflows.

Kulathunga, Dulitha Prasanna [Iowa State Universit↗

mvBayes

SAND2026-16980O mvBayes implements multivariate Bayesian regression using MATLAB and decomposes a multivariate or functional response into components based on a user-specified orthogonal basis. This allows for independent modeling of each component with any chosen univariate Bayesian regression model. This tool includes methods for prediction and visualization, facilitating the evaluation of Bayesian surrogate models through the application of Bayesian theory and Markov Chain Monte Carlo (MCMC) sampling techniques. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗

BayesPPR

SAND2026-17002O BayesPPR performs Bayesian Projection Pursuit Regression (PPR) using MATLAB. A surrogate model for calibration applications, it enables users to efficiently analyze complex datasets and extract meaningful patterns through regression techniques. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗

Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications

Abstract Robust quantification of predictive uncertainty is a critical addition needed for machine learning applied to weather and climate problems to improve the understanding of what is driving prediction sensitivity. Ensembles of machine learning models provide predictive uncertainty estimates in a conceptually simple way but require multiple models for training and prediction, increasing computational cost and latency. Parametric deep learning can estimate uncertainty with one model by predicting the parameters of a probability distribution but does not account for epistemic uncertainty. Evidential deep learning, a technique that extends parametric deep learning to higher-order distributions, can account for both aleatoric and epistemic uncertainties with one model. This study compares the uncertainty derived from evidential neural networks to that obtained from ensembles. Through applications of the classification of winter precipitation type and regression of surface-layer fluxes, we show evidential deep learning models attaining predictive accuracy rivaling standard methods while robustly quantifying both sources of uncertainty. We evaluate the uncertainty in terms of how well the predictions are calibrated and how well the uncertainty correlates with prediction error. Analyses of uncertainty in the context of the inputs reveal sensitivities to underlying meteorological processes, facilitating interpretation of the models. The conceptual simplicity, interpretability, and computational efficiency of evidential neural networks make them highly extensible, offering a promising approach for reliable and practical uncertainty quantification in Earth system science modeling. To encourage broader adoption of evidential deep learning, we have developed a new Python package, Machine Integration and Learning for Earth Systems (MILES) group Generalized Uncertainty for Earth System Science (GUESS) (MILES-GUESS) ( https://github.com/ai2es/miles-guess ), that enables users to train and evaluate both evidential and ensemble deep learning. Significance Statement This study demonstrates a new technique, evidential deep learning, for robust and computationally efficient uncertainty quantification in modeling the Earth system. The method integrates probabilistic principles into deep neural networks, enabling the estimation of both aleatoric uncertainty from noisy data and epistemic uncertainty from model limitations using a single model. Our analyses reveal how decomposing these uncertainties provides valuable insights into reliability, accuracy, and model shortcomings. We show that the approach can rival standard methods in classification and regression tasks within atmospheric science while offering practical advantages such as computational efficiency. With further advances, evidential networks have the potential to enhance risk assessment and decision-making across meteorology by improving uncertainty quantification, a longstanding challenge. This work establishes a strong foundation and motivation for the broader adoption of evidential learning, where properly quantifying uncertainties is critical yet lacking.

Schreck, John S.↗