Search NASA⌕ Search

SEARCH · Search NASA

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Challenges in integrating dissolved organic matter chemodiversity into kinetic models of soil respiration

The chemodiversity of dissolved organic matter (DOM) in soil has been proposed to influence the microbial metabolism and fate of belowground organic carbon (C). However, integrating DOM chemistry into soil C cycle models to improve predictions of C stocks and fluxes—beyond simply considering DOM pool size—remains a challenge. While recent research suggests that incorporating DOM chemodiversity into models can improve predictions of microbial respiration, there is still a lack of mechanistic understanding describing how DOM chemodiversity affects microbial metabolism and soil respiration. Here, we evaluated whether DOM chemodiversity was a determinant of soil respiration using paired measurements of high-resolution DOM chemistry, obtained from Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS), and potential soil respiration rates from across the United States (U.S.), all data provided by the Molecular Observation Network. Our objectives were to (1) assess statistical relationships between DOM chemodiversity and microbial respiration, and (2) evaluate the ability of kinetic models to leverage DOM chemistry to explain empirical relationships found in statistical models. Statistical regressions revealed that DOM chemodiversity (alpha diversity) was nonlinearly related to potential soil respiration rates, both independently and through its interactions with DOM and total C concentrations. In soils with relatively high DOM but low total C concentrations, potential soil respiration rates were negatively correlated with DOM alpha diversity, whereas in soils with relatively low DOM and high total C concentrations showed the opposite trend. However, when metabolic transition theory kinetic models were modified to include chemodiversity, their performance was comparable to traditional Monod kinetics approaches, which simulate respiration rates as a function of DOM concentration. The inability to account for nonlinearities in DOM chemodiversity–respiration relationships highlight an opportunity to advance substrate uptake kinetics by establishing causal links between DOM chemodiversity, microbial metabolism trade-offs, and potential interactions under varied environmental conditions.

Bioenergetic model↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

Search for the associated production of charm quarks and a Higgs boson decaying into a photon pair with the ATLAS detector

A search for the production of a Higgs boson and one or more charm quarks, in which the Higgs boson decays into a photon pair, is presented. This search uses proton-proton collision data with a centre-of-mass energy of $\sqrt{s}$ = 13 TeV and an integrated luminosity of 140 fb −1 recorded by the ATLAS detector at the Large Hadron Collider. The analysis relies on the identification of charm-quark-containing jets, and adopts an approach based on Gaussian process regression to model the non-resonant di-photon background. The observed (expected, assuming the Standard Model signal) upper limit at the 95% confidence level on the cross-section for producing a Higgs boson and at least one charm-quark-containing jet that passes a fiducial selection is found to be 10.6 pb (8.8 pb). The observed (expected) measured cross-section for this process is 5.3 ± 3.2 pb (2.9 ± 3.1 pb).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Temperature and Composition Dependence Modeling of Viscosity and Electrical Conductivity of Low-Activity Waste Glass Melts

The development of models that accurately relate the properties of a glass melt to its temperature and composition is important for glass formulation, melter control, and modeling the melt flow, refractory corrosion, and production rate. Using a database consisting of more than 4,000 data points measured between 900 °C and 1250 °C for over 600 unique low-activity waste glass compositions, we developed models for the melt viscosity and electrical conductivity. Models based on the Gaussian process regression approach outperformed models based on the Vogel–Fulcher–Tammann equation according to four standard metrics and yielded reliable prediction intervals. The models found primarily linear effects between properties and individual components, except for the effect of the Na 2 O mass fraction on the electrical conductivity. The effects were found to be consistent with current theories on physical processes involved with those properties.

36 MATERIALS SCIENCE↗

Protocol to detect dilution cycles in chemostat experiments and estimate growth rate slopes with linear modeling with R software chemostat_regression

Chemostat growth chambers measure optical density over time and require manual calculation of growth rates. Here, we present chemostat_regression, R software that enables users to automatically identify chemostat cycles and estimate growth rate using a linear regression approach. We describe steps for creating requisite software environment(s), formatting input data, executing the software via command line/RStudio/R-Shiny, interpreting results, assessing the validity of results, and modifying input parameters.

59 BASIC BIOLOGICAL SCIENCES↗

Insights into Tetravalent Np Speciation in HNO 3 through Spectroelectrochemistry and Multivariate Analysis

In situ optical spectroscopy, spectropotentiometry, and multivariate analysis were applied to the Np(IV) nitrate system to better understand speciation and quantify HNO 3 concentration. Thin-layer spectropotentiometry, or spectroelectrochemistry, was leveraged to isolate and stabilize Np(IV) without compromising the solution conditions and generate representative Vis-NIR absorption spectra from 0.5 to 10 M HNO 3 and benchmark the corresponding Np(IV) molar absorptivity coefficients. Spectra were described with principal component analysis (PCA) to identify the purest Np(IV) absorbance spectra among other oxidation states [e.g., Np(V/VI)] at each acid concentration and then to identify the primary sources of variance within each Np(IV) spectrum with respect to Np(IV) nitrate complexes. Then, partial least-squares regression (PLSR) and support vector regression (SVR) models were built to predict HNO 3 concentration from the Np(IV) spectral data. The nonlinear SVR model outperformed the linear PLSR model for the HNO 3 concentration predictions. Finally, the inclusion of spectra collected in edge and center point HNO 3 concentrations in the calibration set was determined to be crucial for producing models with strong predictive capabilities. The multivariate approach used in this study makes it possible to quantify HNO 3 concentration solely based on Np(IV) absorption spectra, which is essential to quantifying processing streams in various online monitoring applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Near-Real-Time Material Tracking: Combining Vis–NIR Spectroscopy with Flow Sensing for Accurate Nd(III) Quantification

A fiber-optic visible–near-infrared (vis–NIR) absorption spectroscopy and flow sensor system has been developed for near-real-time tracking of Nd mass in the effluent stream from a column in a fume hood. The approach leverages two unique data streams and a partial least-squares regression (PLSR) model trained on vis–NIR absorption spectra of Nd(III) (0–1.5 M) in 1 M HNO 3 . In-line volumetric flow rate and vis–NIR spectra are measured in sequence after a chromatography column. The time stamps from each data stream are then synchronized, which allows integrated volumes to be combined with Nd(III) molarities predicted by a PLSR model to accurately calculate the Nd mass flowing through the column. This integrated measurement provides instantaneous mass flow and accumulates these data over time to obtain the total mass processed. The methodology developed in this study contributes critical technical infrastructure to improve monitoring capabilities to support chemical separations and the production of strategic materials and isotopes.

Irvine, Sawyer B. [Oak Ridge National Laboratory (↗

BASS

SAND2026-17001O BASS implements Bayesian Adaptive Spline Surfaces in MATLAB and serves as a surrogate model for regression applications. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗

Case Study: Leveraging GenAI to Build AI-based Surrogates and Regressors for Modeling Radio Frequency Heating in Fusion Energy Science

This work presents a detailed case study on using Generative AI (GenAI) to develop AI surrogates for simulation models in fusion energy research. The scope includes the methodology, implementation, and results of using GenAI to assist in model development and optimization, comparing these results with previous manually developed models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Learning the factors controlling mineral dissolution in three-dimensional fracture networks: applications in geologic carbon sequestration

We perform a set of high-fidelity simulations of geochemical reactions within three-dimensional discrete fracture networks (DFN) and use various machine learning techniques to determine the primary factors controlling mineral dissolution. The DFN are partially filled with quartz that gradually dissolves until quasi-steady state conditions are reached. At this point, we measure the quartz remaining in each fracture within the domain as our primary quantity of interest. We observe that a primary sub-network of fractures exists, where the quartz has been fully dissolved out. This reduction in resistance to flow leads to increased flow channelization and reduced solute travel times. However, depending on the DFN topology and the rate of dissolution, we observe substantial variability in the volume of quartz remaining within fractures outside of the primary subnetwork. This variability indicates an interplay between the fracture network structure and geochemical reactions. We characterize the features controlling these processes by developing a machine learning framework to extract their relevant impact. Specifically, we use a combination of high-fidelity simulations with a graph-based approach to study geochemical reactive transport in a complex fracture network to determine the key features that control dissolution. We consider topological, geometric and hydrological features of the fracture network to predict the remaining quartz in quasi-steady state. We found that the dissolution reaction rate constant of quartz and the distance to the primary sub-network in the fracture network are the two most important features controlling the amount of quartz remaining. This study is a first step towards characterizing the parameters that control carbon mineralization using an approach with integrates computational physics and machine learning.

54 ENVIRONMENTAL SCIENCES↗

MINLP for regularized symbolic regression with applications to data-driven modeling of critical minerals processes

The poster summarizes recent advances in symbolic regression developed as part of the PrOMMiS project over the past year. In particular, it describes the comparison of surrogates for critical minerals (CM) & rare earth element (REE) recovery flowsheets obtained via symbolic regression and ALAMO. It also compares the predictive ability and solvability of optimization models that incorporate these surrogates.

36 MATERIALS SCIENCE↗

bayesian_tensor_regression

We plan to release the code used to perform the experiment described in our upcoming publication, entitled “Bayesian Tensor Modeling for Distribution-on-Distribution Regression.” This code a Bayesian regression model with a multi-way Dirichlet prior to tensor input distributions. All code to be released implements a new model that is intended for open-source publication.

Murph, Alexander C. [Los Alamos National Lab]↗

Explaining word embeddings with perfect fidelity: a case study in predicting research impact

The best-performing approaches for scholarly document quality prediction are based on embedding models. In addition to their performance when used in classifiers, embedding models can also provide predictions even for words that were not contained in the labelled training data for the classification model, which is important in the context of the ever-evolving research terminology. Although model-agnostic explanation methods, such as Local interpretable model-agnostic explanations, can be applied to explain machine learning classifiers trained on embedding models, these produce results with questionable correspondence to the model. We introduce a new feature importance method, Self-Model Entities Rated (SMER), for logistic regression-based classification models trained on word embeddings. We show that SMER has theoretically perfect fidelity with the explained model, as the average of logits of SMER scores for individual words (SMER explanation) exactly corresponds to the logit of the prediction of the explained model. Quantitative and qualitative evaluation is performed through five diverse experiments conducted on 50,000 research articles (papers) from the CORD-19 corpus. In conclusion, through an AOPC curve analysis, we experimentally demonstrate that SMER produces better explanations than LIME, SHAP and global tree surrogates.

Coarse-grained models↗

Uncertainty Quantification and Sensitivity Analysis of Non-Nuclear Advanced Controls Testbed Reactor Mockup

The research presented in this report describes our progress in applying stochastic methods and uncertainty quantification, parametric study, and variance-based sensitivity analysis (also known as Sobol sensitivity analysis) to a full-core model of a nuclear thermal propulsion (NTP) system simulated with Griffin, with the goal of developing a reduced order (surrogate) model which can be rapidly sampled while perturbing multiple input parameters. In this NTP system, reactivity and power feedback affect the rotation of control drums, which are controlled by a hybrid proportional, integral and derivative (PID) controller, actuated by the power demand and reactivity feedback from the numerical model. This model uses reactor kinetic feedback (mean generation time and $\beta$ from a transient Griffin simulation executed with the improved quasi-static method to provide the kinetic parameters) as inputs to functions which control the CD rotation angle. Using a number of stochastic method approaches, we developed a dual purpose training-surrogate model of the NTP system using polynomial regression. The trained model can be rapidly sampled while simultaneously perturbing various input parameters of the model, such as coefficients on the PID control, or temperature (directly affect the neutron cross section). The surrogate model delivers accurate results orders-of-magnitude faster (minutes, not days) than the base model. Once the base model has been trained, distributions of the uncertain parameters can be changed at will to investigate the effects of perturbing multiple inputs and their effect on the output. For example, coefficients used in the PID control system may vary due to some physical interference, or there may be uncertainty in the temperature of the neutron cross sections in various regions of the reactor. A distribution can be placed on these parameters and operational boundaries can be determined. The goal of this work is to support development of an advanced control system to operate CDs in a functioning NTP system.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A Tutorial on Bayesian analysis of linear shock compression data

Gas gun and other shock compression experiments often produce shock wave velocity measurements that are linearly associated with particle velocity. Traditionally, this empirical relationship is quantified with a single Hugoniot curve that is estimated using least squares regression. However, for downstream modeling and simulation tasks, it is often more useful to have multiple Hugoniot curves in the pressure–volume plane that are consistent with the data. We employ Bayesian uncertainty quantification methods as a framework for propagating measurement uncertainty through to model parameters and predictions. Specifically, this Tutorial shows how to sample multiple Hugoniot curves in the pressure–volume plane that are consistent with the shock wave-particle velocity measurements in a two-step Bayesian approach. First, we obtain an analytical expression for the posterior distribution of the linear model parameters using Bayesian linear regression. Second, we propagate samples from the posterior distribution through the Rankine–Hugoniot equations to yield Hugoniot curves in the pressure–volume plane. The procedure is demonstrated with publicly available data on argon, copper, and nickel, and compared against bootstrapping and linear regression. The Bayesian procedure is shown to be interpretable, computationally inexpensive, and less sensitive than an alternative bootstrapping approach to the removal of the point in the copper dataset that has the largest particle velocity. As a Tutorial on Bayesian methodology for the shock compression community, we provide several derivations and explanations that make this paper self-contained, and make all code and data available at github.com/llnl/BALSCD.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Data-Driven Digital Twin for Reliability Assessment of DC/DC Buck Converter

In commercial applications, the operation of DC/DC converters significantly impacts overall system performance and long-term reliability. This study introduces a data-driven digital twin (DT) approach for estimating critical degradation parameters of DC/DC BUCK converter under steady-state condition. Initially, a circuit-level MATLAB/Simulink digital model (DM C ) is refined against a hardware prototype’s switching model dataset using offline particle swarm optimization. The optimized digital model’s steady-state response is then verified with its average model response while varying the duty and load. Subsequently, degradation profiles are imposed on the inductor, capacitor, MOSFET in the DMC. A large dataset is generated from this model, allowing training, validation, and testing of machine learning (ML) models for component health regression tasks. The proposed method employs random forest ML models, achieving impressive regression results with a squared R value as high as 0.99978 and a root mean square error of 4.2× 10 –6 . The method is further validated on a medium power level DC/DC BUCK prototype with varying load conditions, and includes the analysis of MOSFET’s on-resistance under degradation conditions. This data-driven DT method shows promise for identifying parasitic degradation and ohmic loss parameters, enhancing converter reliability assessments in a non-invasive, generalized, and computationally efficient manner.

14 SOLAR ENERGY↗

Beyond pinball loss: Quantile methods for calibrated uncertainty quantification

Amongthemanywaysofquantifying uncertainty in a regression setting, specifying the full quantile function is attractive, as quantiles are amenable to interpretation and evaluation. A model that predicts the true conditional quantiles for each input, at all quantile levels, presents a correct and efficient representation of the underlying uncertainty. To achieve this, many current quantile-based methods focus on optimizing the pinball loss. However, this loss restricts the scope of applicable regression models, limits the ability to target many desirable properties (e.g. calibration, sharpness, centered intervals), and may produce poor conditional quantiles. In this work, we develop new quantile methods that address these shortcomings. In particular, we propose methods that can apply to any class of regression model, select an explicit balance between calibration and sharpness, optimize for calibration of centered intervals, and produce more accurate conditional quantiles. We provide a thorough experimental evaluation of our methods, which includes a high dimensional uncertainty quantification task in nuclear fusion.

97 MATHEMATICS AND COMPUTING↗