Search NASASearch

SEARCH · Search NASA

Results for “Gaussian process regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A novel approach for large-scale wind energy potential assessment

Increasing wind energy generation is central to grid decarbonization, yet methods to estimate wind energy potential are not standardized, leading to inconsistencies and even skewed results. This study aims to improve the fidelity of wind energy potential estimates through an approach that integrates geospatial analysis and machine learning (i.e., Gaussian process regression). We demonstrate this approach to assess the spatial distribution of wind energy capacity potential in the Contiguous United States (CONUS). We find that the capacity-based power density ranges from 1.70 MW/km2 (25th percentile) to 3.88 MW/km2 (75th percentile) for existing wind farms in the CONUS. The value is lower in agricultural areas (2.73 ± 0.02 MW/km2, mean ± 95 % confidence interval) and higher in other land cover types (3.30 ± 0.03 MW/km2). Notably, advancements in turbine manufacturing could reduce power density in areas with lower wind speeds by adopting low specific-power turbines, but improve power density in areas with higher wind speeds (>8.35 m/s at 120m above the ground), highlighting opportunities for repowering existing wind farms. Wind energy potential is shaped by wind resource quality and is regionally characterized by land cover and physical conditions, revealing significant capacity potential in the Great Plains and Upper Texas. The results indicate that areas previously identified as hot spots using existing approaches (e.g., the west of the Rocky Mountains) may have a limited capacity potential due to low wind resource quality. Improvements in methodology and capacity potential estimates in this study could serve as a new basis for future energy systems analysis and planning.

Dai, Tao

Direct Deoxygenation of Phenol over Fe-Based Bimetallic Surfaces Using On-the-Fly Surrogate Models

We present an accelerated nudged elastic band (NEB) study of phenol direct deoxygenation (DDO) on Fe-based bimetallic surfaces using a recently developed Gaussian process regression (GPR) calculator. Our test calculations demonstrate that the GPR calculator achieves up to 3 times speedup compared to conventional density functional theory calculations while maintaining high accuracy, with energy barrier errors below 0.015 eV. Using GPR-NEB, we systematically examine the DDO mechanism on pure Fe(110) and surfaces modified with Co and Ni in both top and subsurface layers. Our results show that subsurface Co and Ni substitutions preserve favorable thermodynamics and kinetics for both C–O bond cleavage and C–H bond formation, comparable to those on the pure Fe(110) surface. In contrast, top-layer substitutions generally increase the C–O bond cleavage barrier, render the step endothermic, and result in significantly higher reverse reaction rates, making DDO unfavorable on these surfaces. This work demonstrates the effectiveness of GRR-accelerated transition state searches for complex surface reactions and provides insights into rational design of bimetallic catalysts for selective deoxygenation.

Aromatic compounds

Probing multi-dimensional composition spaces in search of strong metallic alloys

Refractory complex concentrated alloys (RCCA) offer exceptionally high-temperature strength compared to pure metals and dilute alloys, but predictive theory for RCCA design is lacking. We present large-scale molecular Dynamics (MD) simulations of crystal plasticity to explore alloy compositions for maximum mechanical strength, focusing on Fe-Ta-W and Nb-Ta-Mo-W alloy families modeled with Embedded Atom Model (EAM) and Spectral Neighbor Analysis Potentials (SNAP). To efficiently guide the search for strong alloy compositions, we employ iterative optimization using Gaussian process regression. Many simulated RCCA compositions exhibit pronounced cocktail strengthening, with strengths surpassing their strongest constituent metal, tungsten. Contrary to expectations, the highest strength is found on binary edges of the RCCA composition space. Detailed analyses of atomistic simulations reveal that, similar to pure BCC metals, plastic response in RCCA is primarily governed by screw dislocations. However, at large strains, dislocation multiplication and interactions (Taylor hardening) become the dominant mechanisms contributing to RCCA strength.

Materials science

Machine learning inversion from small-angle scattering for charged polymers

We develop Monte Carlo simulations for uniformly charged polymers and a machine learning algorithm to interpret the intra-polymer structure factor of the charged polymer system, which can be obtained from small-angle scattering experiments. The polymer is modeled as a chain of fixed-length bonds, where the connected bonds are subject to bending energy, and there is also a screened Coulomb potential for charge interaction between all joints. The bending energy is determined by the intrinsic bending stiffness, and the charge interaction depends on the interaction strength and screening length. All three contribute to the stiffness of the polymer chain and lead to longer and larger polymer conformations. The screening length also introduces a second length scale for the polymer besides the bending persistence length. To obtain the inverse mapping from the structure factor to these polymer conformation and energy-related parameters, we generate a large data set of structure factors by running simulations for a wide range of polymer energy parameters. We use principal component analysis to investigate the intra-polymer structure factors and determine the feasibility of the inversion using the nearest neighbor distance. We employ Gaussian process regression to achieve the inverse mapping and extract the characteristic parameters of polymers from the structure factor with low relative error.

36 MATERIALS SCIENCE

Machine learning-assisted profiling of a kinked ladder polymer structure using scattering

Ladder polymers consisting of fused rings in the backbone have very limited conformational freedom, which results in very different properties from traditional linear polymers. However, accurately determining their size and chain conformations from solution scattering remains a challenge. Their chain conformations of kinked ladder polymers are largely governed by the structures and relative orientations or configurations of the repeat units, unlike conventional polymer chains whose bending angles between repeat units follow a unimodal Gaussian distribution. Meanwhile, traditional scattering models for polymer chains do not account for these unique structural features. This work introduces a novel approach that integrates machine learning with Monte Carlo simulations to construct a model that can describe the geometry of a type of kinked CANAL ladder polymers. We first develop a Monte Carlo simulation model for sampling the configuration space of CANAL ladder polymers, where each repeat unit is modeled as a biaxial segment. Then, we establish a machine learning-assisted scattering analysis framework based on Gaussian Process Regression. Finally, we conduct small-angle neutron scattering experiments on a CANAL ladder polymer solution to apply our approach. Our method uncovers structural features of such ladder polymers that conventional methods fail to capture.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),

A convergence metric for counting statistics in time-resolved small angle neutron scattering

Here, this work introduces a model-independent, dimensionless metric for predicting optimal measurement duration in time-resolved small-angle neutron scattering using early-time data. Built on a Gaussian process regression framework, the method reconstructs scattering profiles with quantified uncertainty, even from sparse or noisy measurements. Demonstrated on the EQ-SANS instrument at the Spallation Neutron Source, the approach generalizes to general SANS instruments with a two-dimensional detector. A key result is the discovery of a dimensionless convergence metric revealing a universal power-law scaling in profile evolution across soft matter systems. When time is normalized by a system-specific characteristic time t*, the variation in inferred profiles collapses onto a single curve with an exponent between −2 and −1. This trend emerges within the first ten time steps, enabling early prediction of measurement sufficiency. The method supports real-time experimental optimization and is especially valuable for maximizing efficiency in low-flux environments such as compact accelerator-based neutron sources.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN

Bayesian Gaussian process inference for neutron spin echo measurement

Neutron spin echo (NSE) spectroscopy provides unique access to microscopic dynamics, but its application is often constrained by low neutron flux, long acquisition times, and significant noise. Here, we present a Bayesian inference approach based on Gaussian process regression (GPR) to reconstruct high-quality spin echo signals from sparse and noisy data by exploiting correlations in reciprocal space. Benchmarks on synthetic datasets and validation with experimental NSE measurements of dendrimers show that GPR suppresses noise, interpolates missing intensity values, and accommodates irregular observations. The method improves accuracy, shortens acquisition times, and enables high-throughput and real-time studies. Beyond NSE, the framework is broadly applicable to other low signal-to-noise ratio scattering techniques, thereby extending the scope of neutron spectroscopy.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN

Inferring performance metrics for laser direct drive experiments on OMEGA

Quantifying performance improvements on the OMEGA laser facility requires robust inference of established no-alpha performance metrics, which requires, at minimum, a model to infer the shocked fuel mass and pressure of the confined fusion plasma. In this work, we describe the methodology used to infer performance metrics on OMEGA and present the current state-of-the art model used to infer these metrics from OMEGA experiments. In particular, since neutron images of cryogenic implosions are not available on OMEGA at present, we present how x-ray sizes are determined on OMEGA using a Gaussian Process regression model and how the neutron production region's size is inferred from them. As a result, we end by benchmarking the model using synthetic data and 1-D LILAC simulations and test its experimental self-consistency across available x-ray diagnostic channels.

Gopalaswamy, V. [Laboratory for Laser Energetics,

A Bayesian desmearing algorithm for Bonse–Hart USANS with anisotropic scattering

Ultra-small-angle neutron scattering (USANS) using Bonse–Hart optics provides micrometer-scale structural insights but suffers from severe slit-geometry smearing. While well-established for isotropic systems, quantitative desmearing of anisotropic data remains a challenge because conventional corrections break down for non-radial scattering. In this work, we address this by developing a resolution-aware Bayesian framework that explicitly incorporates anisotropy via an affine deformation to the scattering pattern, guided by the principle of parsimony. This results in orientation-resolved point-spread functions that enable a self-consistent determination of both the resolution and deformation parameters. Using Gaussian process regression with uncertainty quantification and a probabilistic correction for multiple scattering, we demonstrate the framework’s effectiveness through numerical benchmarks and experimental studies of a stretched polymer melt. Our approach enables the seamless integration of SANS and USANS data, facilitating quantitative structural analysis of deformed materials at nanometer to micrometer scales.

36 MATERIALS SCIENCE

Optimal binning of correlated measurements

Experimental measurements are commonly represented on a discrete grid, requiring a balance between granularity and statistical noise. Two strategies have traditionally been used to improve such representations: selecting an appropriate bin width to control discretization error and applying kernel-based smoothing to suppress fluctuations. Despite their shared goal, these approaches have largely developed independently, without a unified statistical description of how discretization and correlation jointly determine measurement precision. Here, we extend the discussion of optimal interval averaging to a correlation-aware setting by Gaussian process regression, which explicitly accounts for correlations among neighboring bins. Starting from first principles, we derive the mean-squared error of discretized measurements and obtain closed-form asymptotic expressions for the optimal bin width and correlation length. When recast in reduced variables, the theory reveals distinct universal scaling laws governing the error in the correlation-free and correlation-controlled regimes. Characterized by intrinsically smooth intensity profiles and counting-based statistics, neutron scattering measurements are well suited for demonstrating the enhanced error contraction enabled by inter-bin correlations. We show that such improvement is achievable over the experimentally accessible Q-range and across multiple instruments and material systems. These results show that explicitly accounting for correlations systematically reshapes the limits of precision in discretized, noise-limited measurements. More broadly, the framework provides a transferable statistical foundation for optimizing data representation, inference, and experimental design across the physical and data sciences.

Tung, Chi-Huan [ORNL] (ORCID:0000000221972074)

Aemulus ν: precision halo mass functions in wνCDM cosmologies

Precise and accurate predictions of the halo mass function for cluster mass scales in wνCDM cosmologies are crucial for extracting robust and unbiased cosmological information from upcoming galaxy cluster surveys. Here, we present a halo mass function emulator for cluster mass scales (≳ 1013 M ⊙/h) up to redshift z = 2 with comprehensive support for the parameter space of wνCDM cosmologies allowed by current data. Based on the Aemulus ν suite of simulations, the emulator marks a significant improvement in the precision of halo mass function predictions by incorporating both massive neutrinos and non-standard dark energy equation of state models. This allows for accurate modeling of the cosmology dependence in large-scale structure and galaxy cluster studies. We show that the emulator, designed using Gaussian Process Regression, has negligible theoretical uncertainties compared to dominant sources of error in future cluster abundance studies. Our emulator is publicly available (https://github.com/DelonShen/aemulusnu_hmf), providing the community with a crucial tool for upcoming cosmological surveys such as LSST and Euclid.

cluster counts

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES

Lattice QCD estimates of thermal photon production from the QGP

Thermal photons produced in heavy-ion collision experiments are an important observable for understanding quark-gluon plasma (QGP). The thermal photon rate from the QGP at a given temperature can be calculated from the spectral function of the vector current correlator. Extraction of the spectral function from the lattice correlator is known to be an ill-conditioned problem, as there is no unique solution for a spectral function for a given lattice correlator with statistical errors. The vector current correlator, on the other hand, receives a large ultraviolet contribution from the vacuum, which makes the extraction of the thermal photon rate difficult from this channel. We therefore consider the difference between the transverse and longitudinal part of the spectral function, only capturing the thermal contribution to the current correlator, simplifying the reconstruction significantly. The lattice correlator is calculated for light quarks in quenched QCD at T = 470 MeV ( ∼ 1.5 T c ), as well as in 2 + 1 flavor QCD at T = 220 MeV ( ∼ 1.2 T p c ) with m π = 320 MeV . In order to quantify the nonperturbative effects, the lattice correlator is compared with the corresponding NLO + LPM LO estimate of correlator. The reconstruction of the spectral function is performed in several different frameworks, ranging from physics-informed models of the spectral function to more general models in the Backus-Gilbert method and Gaussian process regression. We find that the resulting photon rates agree within errors. Published by the American Physical Society 2024

Astronomy & Astrophysics

S AP F LOWER : an automated tool for sap flow data preprocessing, gap-filling, and analysis using deep learning

Sap flow, a critical process in plant water use and ecosystem water cycles, is often measured using thermal dissipation probes (TDP) due to their ease of installation and continuous data collection. However, sap flow data frequently include noise, outliers, and gaps, creating challenges for analysis and requiring substantial manual processing. We developed S AP F LOWER , a tool that automates data preprocessing, model training, gap-filling, sapwood area scaling and modeling, and water use analysis. It integrates autocleaning, machine learning and deep learning models (e.g. random forest, Gaussian process regression, long short-term memory (LSTM), bidirectional LSTM (BiLSTM)), and efficient workflows to process sap flow data. S AP F LOWER can remove over 90% of noisy data while preserving legitimate variations and achieve high accuracy in gap-filling based on user-determined parameters. Random forest, LSTM, and BiLSTM models reduced root mean square error to 10% or less for long-term gaps. Model training and prediction can be performed efficiently within seconds. S AP F LOWER significantly enhances the efficiency and accessibility of TDP data analysis by automating complex tasks, enabling researchers without programming expertise to employ advanced techniques. Future improvements will focus on species-specific corrections for TDP and support for additional measurement methods. S AP F LOWER is openly available on GitHub (https://github.com/JiaxinWang123/SapFlower) and Zenodo (doi: 10.5281/zenodo.13665919).

ecosystem water balance

GP-BayesOpInf

SAND2025-01851O GP-BayesOpInf is a software tool that uses algorithms to combine Gaussian process regression, principal component analysis, and linear Bayesian inference to produce a probabilistic reduced-order model for time-dependent systems. Numerical examples include the compressible Euler equations for an ideal gas, a heat diffusion process with a nonlinear reaction term, and a set of ordinary differential equations describing a compartmental model in epidemiology. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC

Machine Learning Approach for Spatiotemporal Multivariate Optimization of Environmental Monitoring Sensor Locations

Abstract Long-term environmental monitoring is critical for managing the soil and groundwater at contaminated sites. Recent improvements in state-of-the-art sensor technology, communication networks, and artificial intelligence have created opportunities to modernize this monitoring activity for automated, fast, robust, and predictive monitoring. In such modernization, it is required that sensor locations be optimized to capture the spatiotemporal dynamics of all monitoring variables as well as to make it cost-effective. The legacy monitoring datasets of the target area are important to perform this optimization. In this study, we have developed a machine-learning approach to optimize sensor locations for soil and groundwater monitoring based on ensemble supervised learning and majority voting. For spatial optimization, Gaussian process regression (GPR) is used for spatial interpolation, while the majority voting is applied to accommodate the multivariate temporal dimension. Results show that the algorithms significantly outperform the random selection of the sensor locations for predictive spatiotemporal interpolation. While the method has been applied to a four-dimensional dataset (with two-dimensional space, time, and multiple contaminants), we anticipate that it can be generalizable to higher-dimensional datasets for environmental monitoring sensor location optimization.

Siddiquee, Masudur R.

Extratropical Cloud Feedback Constrained by Cloud Sources and Sinks in Cyclones

Constraining cloud feedback in global climate models (GCMs) using observations is important for establishing accurate predictions of future climate. Uncertainty in shortwave cloud feedback (SW FB ) dominates uncertainty in total cloud feedback. Recent studies show a shift toward more positive extratropical SW FB in the latest generations of GCMs leading to the emergence of very high equilibrium climate sensitivity (ECS). In this study, we use precipitation efficiency and albedo susceptibility to constrain liquid water path (LWP) response to warming and SW FB in the Southern Ocean (SO; 50°–80°S). We analyze precipitation in extratropical cyclones (ECs) to learn about extratropical condensed water sink processes, combined with observations of clouds and moisture convergence, and use the analysis to better understand and constrain SW FB . We utilize a perturbed parameter ensemble (PPE) hosted in the Community Atmosphere Model, version 6 (CAM6), to provide a constraint on SW FB based on observations from Clouds and the Earth’s Radiant Energy System (CERES) and Multisensor Advanced Climatology of LWP (MAC-LWP). We apply Gaussian process regression to emulate the model response to all parameters perturbed in the PPE. Confronting the emulator output with observations provides a new estimated response of Earth to global warming. Furthermore, our new estimates of SO LWP reduce the PPE range by 66%–72%, which results in a shortwave cloud radiative effect estimated range that is 27%–34% less than the PPE range. Observations suggest a more positive SO SW FB than the Community Earth System Model, version 2 (CESM2), and consequently do not reject the high climate sensitivity GCMs emerging from the Coupled Model Intercomparison Project phase 6 (CMIP6).

Atmosphere

Machine learning enhanced characterization and optimization of photonic cured MAPbI 3 for efficient perovskite solar cells

Photonic curing (PC) can facilitate high-speed perovskite solar cell (PSC) manufacturing because it uses high-intensity light pulses to crystallize perovskite films in milliseconds. However, optimizing PC conditions is challenging due to its many variables, and using power conversion efficiency (PCE) as the optimization metric is both time-consuming and labor-intensive. This work presents a machine learning (ML) approach to optimize PC conditions for fabricating methylammonium lead iodide (MAPbI 3 ) films by quantitatively comparing their ultraviolet-visible (UV-vis) absorbance spectra to thermal annealed (TA) films using four similarity metrics. We perform Bayesian optimization coupled with Gaussian process regression (BO-GP) to minimize the similarity metrics. Refining PC conditions using active learning based on BO-GP models, we achieve a PC MAPbI3 film with an absorbance spectrum closely matching a TA reference film, which is further verified by its crystalline and morphological properties. Thus, we demonstrate that the UV-vis absorption spectrum can accurately proxy film quality. Additionally, we use an AI-based segmentation model for a more efficient grain size analysis. However, when we use the optimized PC condition to fabricate PSCs, we find that interaction between MAPbI 3 and the hole transport layer (HTL) during PC critically degrades the PSC performance. By adding a buffer layer between the HTL and MAPbI 3 , the optimized PC PSCs produce a champion PCE of 11.8%, comparable to the TA reference of 11.7%. Using UV-vis similarity metrics instead of device PCE as the objective in our BO-GP method accelerates the optimization of PC processing conditions for MAPbI 3 films.

14 SOLAR ENERGY