Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Beyond Point Estimates: Benchmarking Uncertainty Quantification Methods on the AION-1 Astronomical Foundation Model

Foundation models for astronomical surveys offer powerful learned representations that can be transferred to downstream regression tasks such as galaxy property estimation. However, point predictions alone are insufficient for scientific inference; reliable uncertainty quantification (UQ) is essential. We compare seven UQ methods on galaxy property regression using frozen AION-1 foundation-model embeddings, predicting redshift, stellar mass, stellar-population age, gas-phase metallicity, and specific star-formation rate, from Legacy Survey photometry/imaging and DESI spectra, with PROVABGS-derived labels. Distribution-free conformal methods achieve marginal coverage within $\sim$1 pp of the nominal 90% across all properties, while non-conformal baselines (Deep Ensembles, MC~Dropout) fail to calibrate reliably. Among conformal approaches, Conformalized Quantile Regression (CQR) delivers the best coverage in the bin with the poorest model predictions. More importantly, only the Locally Valid and Discriminative (LVD) framework -- particularly when operating on AION-1 embeddings -- also provides finite-sample \emph{local validity}, producing intervals that adapt to each galaxy's local prediction difficulty rather than relying on marginal guarantees alone. These results establish conformal prediction, and LVD in particular, as the preferred UQ framework for uncertainty-aware inference on foundation-model embeddings in astrophysics.

Tame-Narvaez, Karla [Fermilab] (ORCID:000000022249↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Rate expressions and kinetic parameters for metal ferrites in relation to applications of fossil fuel conversion to hydrogen: Part 1 of 2

Here, the goal of the present work was to provide the necessary reaction emulation information to enable detailed process simulation of a chemical looping H 2 production system from fossil fuels using CaFe 2 O 4 . This specifically pertained to the necessary kinetic data, reaction model development, and model rate parameters required for reaction emulation in both reducing and oxidizing environments. A logical methodology was defined, which included discretization of the reaction network, establishing a core model for reaction emulation that could be adapted based on the system phenomena, and development of a rate parameter regression tool designed around the core model. An extensive array of data sets was acquired by which parametric regressions were performed. The work presented and tabulated a comprehensive set of rate parameters for the reduction and oxidation reactions of CaFe 2 O 4 and descendent phases of Ca 2 Fe 2 O 5 , FeO, Fe 3 O 4 , Fe, and CaO to emulate reaction behavior in a looping-based process environment. This included direct reduction using CH 4 , H 2 , and CO, and direct oxidation reactions with steam, CO 2 and O 2 . Dynamic equilibrium was quantified for reactions that could utilize H 2 O and CO 2 as soft oxidants to re-saturate lattice oxygen in the depleted structure/phases. The kinetics associated with the oxidative mechanisms with the soft oxidants were quantified and compared to those of the reducing counterparts. The analysis provided critical insight to emulate reactions for a process that seeks to use natural gas (NG) or other fossil fuels as a direct reductant for the end goal of H 2 production.

calcium ferrite oxygen carriers↗

Short-Term Load Forecasting Considering EV Charging Loads with Prediction Interval Evaluation

Short-term load forecasting plays a critical role in power system planning and operation. Along with the electrification of various loads, electricity demands are becoming increasingly hard to predict. Notably, the recent rise in electric vehicles (EVs) has further contributed to this unpredictability. To address this issue, this paper proposes a probabilistic load forecasting strategy utilizing Gaussian process regression, structured in a day-ahead manner. While many works focus on deterministic prediction, probabilistic forecasting offers additional insights into variability and uncertainty, enabling more flexible and reliable operation for power systems. To enhance the accuracy of the load forecasting model, the inputs include features related to EV charging habits as well as commonly used weather information. The load forecasting results are evaluated using various metrics, including conventional ones that assess the accuracy of point forecasts, as well as additional metrics that test the reliability of prediction intervals. The proposed load forecasting method is finally tested on real residential power consumption data and EV charging data sampled from real-world sources. The results prove that the new features can greatly improve the performance of the load forecasting method.

electrical vehicle↗

Joint Modeling of Wind Speed and Wind Direction Through a Conditional Approach

Atmospheric near surface wind speed and wind direction play an important role in many applications, ranging from air quality modeling, building design, wind turbine placement to climate change research. It is therefore crucial to accurately estimate the joint probability distribution of wind speed and direction. In this work, we develop a conditional approach to model these two variables, where the joint distribution is decomposed into the product of the marginal distribution of wind direction and the conditional distribution of wind speed given wind direction. To accommodate the circular nature of wind direction, a von Mises mixture model is used; the conditional wind speed distribution is modeled as a directional dependent Weibull distribution via a two-stage estimation procedure, consisting of a directional binned Weibull parameter estimation, followed by a harmonic regression to estimate the dependence of the Weibull parameters on wind direction. A Monte Carlo simulation study indicates that our method outperforms two other approaches in estimation efficiency: one that utilizes periodic spline quantile regression and another that generates data from the commonly used Abe-Ley distribution for cylindrical data. We illustrate our method by using the output from a regional climate model to investigate how the joint distribution of wind speed and direction may change under some future climate scenarios. Our method indicates significant changes in the variation of wind speed with respect to some directions.

17 WIND ENERGY↗

Computationally efficient and error aware surrogate construction for numerical solutions of subsurface flow through porous media

Limiting the injection rate to restrict the pressure below a threshold at a critical location can be an important goal of simulations that model the subsurface pressure between injection and extraction wells. The pressure is approximated by the solution of Darcy’s partial differential equation for a given permeability field. The subsurface permeability is modeled as a random field since it is known only up to statistical properties. This induces uncertainty in the computed pressure. Solving the partial differential equation for an ensemble of random permeability simulations enables estimating a probability distribution for the pressure at the critical location. These simulations are computationally expensive, and practitioners often need rapid online guidance for real-time pressure management. An ensemble of numerical partial differential equation solutions is used to construct a Gaussian process regression model that can quickly predict the pressure at the critical location as a function of the extraction rate and permeability realization. The Gaussian process surrogate analyzes the ensemble of numerical pressure solutions at the critical location as noisy observations of the true pressure solution, enabling robust inference using the conditional Gaussian process distribution. Our first novel contribution is to identify a sampling methodology for the random environment and matching kernel technology for which fitting the Gaussian process regression model scales as O ( n log n ) instead of the typical O ( n 3 ) rate in the number of samples n used to fit the surrogate. The surrogate model allows almost instantaneous predictions for the pressure at the critical location as a function of the extraction rate and permeability realization. Our second contribution is a novel algorithm to calibrate the uncertainty in the surrogate model to the discrepancy between the true pressure solution of Darcy’s equation and the numerical solution. Finally, although our method is derived for building a surrogate for the solution of Darcy’s equation with a random permeability field, the framework broadly applies to solutions of other partial differential equations with random coefficients.

54 ENVIRONMENTAL SCIENCES↗

Data driven investigation to understand the influence of total solids on biological biogas upgrading

In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.

In situ biogas upgrading↗

Modeling plasticity-mediated void growth at the single crystal scale: A physics-informed machine learning approach

Modeling the evolution of voids during plastic flow as well as their effects on plastic dissipation is critical for both component manufacturing and lifetime estimation purposes. To this end, we propose a rate-dependent constitutive model to homogenize the effects of semi-randomly distributed voids on single crystal plasticity whilst capturing void interaction and plastic anisotropy. Here, this present work focuses on the case of face centered cubic crystals to introduce an anisotropic gauge function applicable within the crystal plasticity formalism. The approach combines analytical methods to describe the micromechanics of the system in combination with symbolic regression to capture analytically intractable mechanisms from data. The hybrid framework uses a physics-informed genetic programming-based symbolic regression algorithm to solve a multiform optimization problem simultaneously producing a new gauge function and a new strain rate equation. This is also a multi-objective optimization problem with many competing objectives. A new search and selection step is introduced to the genetic algorithm that promotes convergence toward a global solution that better satisfies all the objectives. Overall, the symbolic equations produced leverage data-driven methods to achieve greater accuracy than comparable alternatives on an analytically intractable problem while maintaining model transparency.

36 MATERIALS SCIENCE↗

Predicting Damages to Remainder Parcels in Right-of-Way Acquisitions for Expanding Transportation Infrastructure: Using a Truncated Finite-Mixture Model

Right-of-way acquisition is a critical component of transportation infrastructure development. Transportation infrastructure projects cannot proceed without proper right-of-way acquisition or may face significant delays. State Departments of Transportation frequently acquire parcels of land for roadway expansion projects. A majority of these acquisitions can be partial takings, referring to a portion of a parcel that is acquired. The remainder of the property usually suffers economic changes due to the partial acquisition, which can be calculated as damage percentages. The damage percentage represents the extent to which the remaining land or property value has been diminished due to the acquisition. It reflects the remaining property value percentage that may have been lost or compromised due to the acquisition. Here, this study aims to provide a robust model to estimate damage percentages to the remainder parcels that may help state Departments of Transportation appraisers make early predictions about the damages in cases involving partial takings. The research uses 509 appraisal reports from the Tennessee Department of Transportation to identify the key parcel attributes that influence the percentage of damages. Three regression models are developed: a linear regression model, a finite-mixture model (FMM), and a truncated FMM with two latent classes. The modeling results show that the truncated FMM with two classes outperforms the other models. To validate the models, actual sales data is collected and analyzed for 59 properties, and the results suggest that the model predictions are fairly accurate. A predictive tool is developed based on the models to help appraisers anticipate right-of-way damages under different scenarios and can provide early predictions about the damages.

42 ENGINEERING↗

Shock Hugoniot calculations using on-the-fly machine learned force fields with ab initio accuracy

We present a framework for computing the shock Hugoniot using on-the-fly machine learned force field (MLFF) molecular dynamics simulations. In particular, we employ an MLFF model based on the kernel method and Bayesian linear regression to compute the free energy, atomic forces, and pressure, in conjunction with a linear regression model between the internal and free energies to compute the internal energy, with all training data generated from Kohn–Sham density functional theory (DFT). We verify the accuracy of the formalism by comparing the Hugoniot for carbon with recent Kohn–Sham DFT results in the literature. In so doing, we demonstrate that Kohn–Sham calculations for the Hugoniot can be accelerated by up to two orders of magnitude, while retaining ab initio accuracy. We apply this framework to calculate the Hugoniots of 14 materials in the FPEOS database, comprising 9 single elements and 5 compounds, between temperatures of 10 kK and 2 MK. We find good agreement with first principles results in the literature while providing tighter error bars. In addition, we confirm that the inter-element interaction in compounds decreases with temperature.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE↗

Unifying simulation and inference with normalizing flows

There have been many applications of deep neural networks to detector calibrations and a growing number of studies that propose deep generative models as automated fast detector simulators. We show that these two tasks can be unified by using maximum likelihood estimation (MLE) from conditional generative models for energy regression. Unlike direct regression techniques, the MLE approach is prior independent and non-Gaussian resolutions can be determined from the shape of the likelihood near the maximum. Using an ATLAS-like calorimeter simulation, we demonstrate this concept in the context of calorimeter energy calibration. Published by the American Physical Society 2025

Hadronic calorimiters↗

Taylor-Expansion-Based Robust Power Flow in Unbalanced Distribution Systems: A Hybrid Data-Aided Method

Traditional power flow methods often adopt certain assumptions designed for passive balanced distribution systems, thus lacking practicality for unbalanced operation. moreover, their computation accuracy and efficiency are heavily subject to unknown errors and bad data in measurements or prediction data of distributed energy resources (ders). to address these issues, this paper proposes a hybrid data-aided robust power flow algorithm in unbalanced distribution systems, which combines taylor series expansion knowledge with a data-driven regression technique. the proposed method initiates a linearization power flow model to derive an explicitly analytical solution by modified taylor expansion. to mitigate the approximation loss that surges due to the der integration and bad data, we further develop a data-aided robust support vector regression approach to estimate the errors efficiently. comparative analysis in the 13-bus and 123-bus ieee unbalanced feeders shows that the proposed hybrid algorithm achieves superior computational efficiency, with guaranteed accuracy and robustness against outliers.

data-driven↗

The Effect of Updraft Entrainment on Convective Cell Deepening in Realistic Large-Eddy Simulations

Entrainment of surrounding cooler and drier air into convective updrafts is one of the key processes that influence deep convection initiation and growth. Numerous studies have investigated the effect of entrainment on isolated convective cloud growth in idealized simulations, but the importance of this effect in realistic conditions with many interacting convective clouds remains uncertain. We examine the impact of entrainment on the depth reached by convective clouds in realistic large-eddy simulations (LES) over central Argentina during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign. Cloudy updrafts and their associated properties are assigned to convective cells tracked with radar reflectivity signatures. Several thousand convective cells are tracked over two high convective available potential energy (CAPE) and two low CAPE cases that support cells of varying depths and intensities. Entrainment is calculated explicitly as the fluxes of air into the outer surface of each cloudy updraft. Single-predictor logistic regression models are used to determine the relative importance of updraft, near-updraft, and preconvective initiation atmospheric conditions in predicting whether convective cells become deep. We then build a multiple-predictor regression model pairing important updraft and meteorological metrics with fractional entrainment rate. The probability of cells transitioning to deep convection is most sensitive to ambient 600-hPa relative humidity (42% of total metric contribution to cloud depth predictability), followed by low-level CAPE (28%), cloud-base updraft width (19%), and fractional entrainment (11%). Thus, the initial width of the updraft along with potential buoyancy and its dilution through the midtroposphere collectively determine whether deep convection will result from shallower clouds.

54 ENVIRONMENTAL SCIENCES↗

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Corrected Score Function Framework for Modelling Circadian Gene Expression

Many biological processes display oscillatory behaviour based on an approximately 24 h internal timing system specific to each individual. One process of particular interest is gene expression, for which several circadian transcriptomic studies have identified associations between gene expression during a 24 h period and an individual's health. A challenge with analysing data from these studies is that each individual's internal timing system is offset relative to the 24 h day-night cycle, where day–night cycle time is recorded for each collected sample. Laboratory procedures can accurately determine each individual's offset and determine the internal time of sample collection. However, these laboratory procedures are labour-intensive and expensive. Here, in this paper, we propose a corrected score function framework to obtain a regression model of gene expression given internal time when the offset of each individual is too burdensome to determine. A feature of this framework is that it does not require the probability distribution generating offsets to be symmetric with a mean of zero. Simulation studies validate the use of this corrected score function framework for cosinor regression, which is prevalent in circadian transcriptomic studies. Illustrations with data from three circadian transcriptomic studies further demonstrate that the proposed framework consistently mitigates bias relative to using a score function that does not account for this offset.

59 BASIC BIOLOGICAL SCIENCES↗

A Data-Driven Method for Modeling Creep-Fatigue Stress- Strain Behavior Using Neural ODEs

In this paper, we introduce a data-driven machine learning approach for modeling one-dimensional stress–strain behavior under cyclic loading, utilizing experimental data from the nickel-based Alloy 617. The study employs uniaxial creep–fatigue test data acquired under various loading histories and compares two distinct neural network-based ODE models. The first model, known as the black-box model, comprehensively describes the strain–stress relationship using a Neural ODE equation. To interpret this black-box model, we apply the Sparse Identification of Nonlinear Dynamical Systems (SINDy) technique, transforming the black-box model into an equation-based model using symbolic regression. The second model, the Neural flow rule model, incorporates Hooke’s Law for the linear elastic component, with the nonlinear part characterized by a Neural ODE. Both models are trained with experimental data to accurately reflect the observed stress–strain behavior. We conduct a detailed comparison with the standard Chaboche model, which includes three back stresses. Our results demonstrate that the neural network-based ODE models precisely capture the experimental creep–fatigue mechanical behavior, exceeding the standard Chaboche model’s accuracy. Furthermore, an interpretable model derived from the black-box neural ODE model through symbolic regression achieves accuracy comparable to the Chaboche model, enhancing its interpretability. The results highlight the potential of neural network-based ODE models to depict complex creep–fatigue behavior, eliminating the necessity for experts to define a specific, material-focused model form.

creep-fatigue↗

Temperature-dependent mechanical properties and crystal plasticity parameters for additively manufactured Haynes-214 alloy: Experiments and numerical modeling

Our experimental mechanical testing data demonstrated that the additively manufactured (AM) laser powder bed fusion (L-PBF) Haynes-214 alloy exhibits non-linear mechanical properties as the temperature rises from ambient to 870 °C. Crystal plasticity (CP) simulations provide an effective approach to gaining deeper insights into microstructure-property linkages under thermomechanical loading. This method can reduce the need for costly high-temperature mechanical testing while accounting for the effects of crystallographic texture and grain morphology on the mechanical behavior of AM materials. However, calibrating a CP model is time-consuming because individual simulations are computationally expensive and hundreds (or more) of iterations over parameter sets may be required. To address this issue, we have designed a machine learning-differential evolution (ML-DE) CP framework that can accurately interpolate the tensile properties of AM L-PBF Haynes-214 alloy across a wide temperature range from ambient to 870 °C, with minimal reliance on experimental data. The framework uses electron backscatter diffraction (EBSD) measurements to generate statistically equivalent microstructural volume elements to serve as inputs to the CP modeling framework. Stress–strain curves were generated from 1000 CP simulations, which serve as the training data set for the three ML regression algorithms explored: linear, extra-trees, and multi-layer perceptron. These three regression models were independently evaluated to compare their efficiency and identify the most suitable algorithm for the given problem. Results revealed that the extra-trees ML regressor outperforms the other models in both qualitative and quantitative aspects with an R 2 of 0.98. Subsequently, the differential evolution optimization approach is employed to calibrate the ML-based CP material parameters with experimental results obtained at various temperatures. Finally, temperature-dependent CP material parameters are formulated. The effectiveness and efficiency of the designed framework are validated through comparison with experimental results, demonstrating a high degree of agreement. These calibrated parametric constitutive equations enable further use of the CP model to study the deformation behavior of this alloy under a wide range of thermo-mechanical loading conditions.

36 MATERIALS SCIENCE↗