Search NASASearch

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Binding energy of the 𝑇 𝑏⁢𝑏 tetraquark from lattice QCD with relativistic and nonrelativistic heavy-quark actions

We present a new determination of the $b\bar{b}$𝑢⁢𝑑 (𝐽 𝑃 = 1 + , 𝐼 = 0) tetraquark binding energy using lattice quantum chromodynamics (QCD) with domain-wall light quarks and a nonperturbatively tuned three-parameter anisotropic-clover “relativistic” action for the 𝑏 quarks. We also perform a direct comparison with a reanalysis of data generated in prior work using a lattice-nonrelativistic QCD (NRQCD) action for the 𝑏 quarks and otherwise identical parameters. Using the new data with relativistic 𝑏 quarks from seven different ensembles with multiple lattice spacings and pion masses, we perform combined chiral and continuum extrapolations and obtain (𝑚 𝑇 𝑏⁢𝑏 −𝑚 𝐵 −𝑚 𝐵* ) RHQ =(−76 ±23) MeV. For the NRQCD data from five ensembles, we perform chiral-only extrapolations and obtain (𝑚 𝑇 𝑏⁢𝑏 −𝑚 𝐵 −𝑚 𝐵* ) NRQCD = (−74 ±17 ±10) MeV. The lower magnitude of the results obtained here, compared to the original analysis in [Phys. Rev. D 100, 014503 (2019)], is due to the use of the symmetric parts of the correlation matrices with local four-quark operators only.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Advanced Semi-Supervised Learning with Uncertainty Estimation for Phase Identification in Distribution Systems

The integration of advanced metering infrastructure (AMI) into power distribution networks generates valuable data for tasks such as phase identification; however, the limited and unreliable availability of labeled data in the form of customer phase connectivity presents challenges. To address this issue, we propose a semi-supervised learning (SSL) framework that effectively leverages labeled and unlabeled data. Our approach incorporates self-training, label spreading, and Bayesian neural networks (BNNs) to enhance phase identification with AMI data. Our method uses an ensemble of multilayer perceptron classifiers in a self-training setup, iteratively adding high-confidence pseudo-labels to improve robustness. We also apply label spread to propagate labels based on data similarity, which enhances generalization across diverse distributions. In addition, we employ a BNNs with uncertainty estimation, boosting confidence in predictions and reducing phase identification errors. In our case study, we achieved approximately 98% +/- 0.08 accuracy with uncertainty using minimal and unreliable labeled data from a real U.S. utility, Duquesne Light Company. Our SSL approach, combined with uncertainty estimation, provides an efficient solution for phase identification in AMI data, ultimately improving the reliability of smart grid applications.

24 POWER TRANSMISSION AND DISTRIBUTION

Advanced Semi-Supervised Learning With Uncertainty Estimation for Phase Identification in Distribution Systems

The integration of advanced metering infrastructure (AMI) into power distribution networks generates valuable data for tasks such as phase identification; however, the limited and unreliable availability of labeled data in the form of customer phase connectivity presents challenges. To address this issue, we propose a semi-supervised learning (SSL) framework that effectively leverages labeled and unlabeled data. Our approach incorporates self-training, label spreading, and Bayesian neural networks (BNNs) to enhance phase identification with AMI data. Our method uses an ensemble of multilayer perceptron classifiers in a self-training setup, iteratively adding high-confidence pseudo-labels to improve robustness. We also apply label spread to propagate labels based on data similarity, which enhances generalization across diverse distributions. In addition, we employ a BNNs with uncertainty estimation, boosting confidence in predictions and reducing phase identification errors. In our case study, we achieved approximately 98% +/- 0.08 accuracy with uncertainty using minimal and unreliable labeled data from a real U.S. utility, Duquesne Light Company. Our SSL approach, combined with uncertainty estimation, provides an efficient solution for phase identification in AMI data, ultimately improving the reliability of smart grid applications.

24 POWER TRANSMISSION AND DISTRIBUTION

A Practical Probabilistic Benchmark for AI Weather Models

Since the weather is chaotic, it is necessary to forecast an ensemble of future states. Recently, multiple AI weather models have emerged claiming breakthroughs in deterministic skill. Unfortunately, it is hard to fairly compare ensembles of AI forecasts because variations in ensembling methodology become confounding and the baseline data volume is immense. We address this by scoring lagged initial condition ensembles—whereby an ensemble can be constructed from a library of deterministic hindcasts. This allows the first parameter‐free intercomparison of leading AI weather models' probabilistic skill against an operational baseline. Lagged ensembles of the two leading AI weather models, GraphCast and Pangu, perform similarly even though the former outperforms the latter in deterministic scoring. These results are elaborated upon by sensitivity tests showing that commonly used multiple time‐step loss functions damage ensemble calibration.

54 ENVIRONMENTAL SCIENCES

Transfer learning of neural surrogates on multifidelity groundwater simulations

Multifidelity data used in the paper published in Advances in Water Resources 206 (2025) 105140, https://doi.org/10.1016/j.advwatres.2025.105140 The code used to process the data is openly available on GitHub at https://github.com/Model-Reduction-and-UQ-Group/Transfer_Learning_K_reconstruction Computationally inexpensive surrogates of process-based models, such as deep neural networks, enable ensemble-based computations used in risk assessment, data assimilation, etc. However, generation of large datasets required to train a neural network can be as expensive as the ensemble simulations themselves. We ameliorate this challenge by using data from multifidelity (MF) groundwater simulations and transfer learning (TL) to reduce data generation costs while maintaining model accuracy. As a computational example, we train a deep convolutional neural network (CNN) to reconstruct permeability fields from saturation maps derived from a multiphase flow model. Starting with very low- and low-fidelity data generated on increasingly coarse meshes, we pretrain the CNN, followed by output-layer training and fine-tuning using only a limited number of high-fidelity samples. We demonstrate the surrogate’s robustness when interpreting low-quality inputs—such as interpolated maps or data affected by noise—which has strong implications for the applicability in practical hydrogeological scenarios. This multilevel MF-TL strategy achieves a favorable trade-off between computational efficiency and predictive accuracy, significantly outperforming high-fidelity-only approaches under the same computational budget.

Chiofalo, Alessia [University of Bologna] (ORCID:0

From Points to Planes: A Workflow for Converting Three‐Dimensional Point Cloud Data Into Discrete Fracture Network Flow and Transport Models

We present the Point cLoud Algorithm for NEtwork Extraction of Discrete Fracture Networks (PLANE-DFN), a point cloud–based algorithm for automatic fracture network extraction designed to support discrete fracture network (DFN) modeling workflows. PLANE-DFN segments three-dimensional fracture planes from raw point cloud data using RANdom SAmple Consensus coupled with statistical outlier removal and density-based clustering to isolate individual fracture features. Each candidate plane is constrained against site-specific structural constraints based on strike and dip. After segmentation, each fracture is converted into a 2-D convex polygon suitable for meshing and simulation. The PLANE-DFN algorithm is validated by comparing geometric and flow and transport data against data from dfnWorks simulations with ensembles of plane-fit networks. We find that the flow and transport in plane-fit networks are comparable to dfnWorks-generated networks when realistic network geometry is maintained. The PLANE-DFN algorithm provides an automated and streamlined workflow to transform point clouds of data into DFN network geometry.

54 ENVIRONMENTAL SCIENCES

Nonlinear Ensemble Filtering with Diffusion Models: Application to the Surface Quasigeostrophic Dynamics

The intersection between classical data assimilation methods and novel machine learning techniques has attracted significant interest in recent years. Here, we explore another promising solution in which diffusion models are used to formulate a robust nonlinear ensemble filter for sequential data assimilation. Unlike standard machine learning methods, the proposed ensemble score filter (EnSF) is completely training free and can efficiently generate a set of analysis ensemble members. Here, in this study, we apply the EnSF to a surface quasigeostrophic model and compare its performance against the popular local ensemble transform Kalman filter (LETKF), which makes Gaussian assumptions in the analysis step. Numerical tests demonstrate that EnSF maintains stable performance in the absence of localization and for a variety of experimental settings. We find that while LETKF maintains optimal performance in the case of linear observations of the entire state and a perfect model, EnSF shows improvements over LETKF when nonlinear observations are assimilated and the system is subject to unexpected model errors. A spectral decomposition of the analysis results in this nonlinear observation regime shows that the largest improvements over LETKF occur at large scales (small wavenumbers), where LETKF lacks sufficient ensemble spread. Overall, this initial application of EnSF to a geophysical model of intermediate complexity motivates further development of the algorithm for more realistic problems.

Artificial intelligence

Exascale Computing and Data Handling: Challenges and Opportunities for Weather and Climate Prediction

The emergence of exascale computing and artificial intelligence offer tremendous potential to significantly advance Earth system prediction capabilities. However, enormous challenges must be overcome to adapt models and prediction systems to use these new technologies effectively. A 2022 WMO report on exascale computing recommends “urgency in dedicating efforts and attention to disruptions associated with evolving computing technologies that will be increasingly difficult to overcome, threatening continued advancements in weather and climate prediction capabilities.” Further, the explosive growth in data from observations, model and ensemble output, and postprocessing threatens to overwhelm the ability to deliver timely, accurate, and precise information needed for decision-making. Artificial intelligence (AI) offers untapped opportunities to alter how models are developed, observations are processed, and predictions are analyzed and extracted for decision-making. Given the extraordinarily high cost of computing, growing complexity of prediction systems, and increasingly unmanageable amount of data being produced and consumed, these challenges are rapidly becoming too large for any single institution or country to handle. This paper describes key technical and budgetary challenges, identifies gaps and ways to address them, and makes a number of recommendations.

Atmosphere

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES

Hyperspectral segmentation of plants in fabricated ecosystems

Hyperspectral imaging provides a powerful tool for analyzing above-ground plant characteristics in fabricated ecosystems, offering rich spectral information across diverse wavelengths. This study presents an efficient workflow for hyperspectral data segmentation and subsequent data analytics, minimizing the need for user annotation through the use of ensembles of sparse mixed scale convolution neural networks. The segmentation process leverages the diversity of ensembles to achieve high accuracy with minimal labeled data, reducing labor-intensive annotation efforts. To further enhance robustness, we incorporate image alignment techniques to address spatial variability in the dataset. Downstream analysis focuses on using the segmented data for processing spectral data, enabling monitoring of plant health. This approach provides a scalable solution for spectral segmentation, and facilitates actionable insights into plant conditions in complex, controlled environments. Our results demonstrate the utility of combining advanced machine learning techniques with hyperspectral analytics for high-throughput plant monitoring.

Zwart, Petrus H.

Contrasting Parametric Sensitivities in Two Global Vegetation Models Using Parameter Perturbation Ensembles

Uncertainty in land model projections remains high and the roles of parametric and structural uncertainty are difficult to disentangle. To compare parametric sensitivity across model structures we present two parameter perturbation ensembles using the Community Land Model (CLM) operating in satellite phenology mode. The ensembles contrast two vegetation modules: (a) the default CLM vegetation module and (b) the Functionally Assembled Terrestrial Ecosystem Simulator (CLM-FATES). We perturbed over 300 parameters and quantified their effects on biophysical fluxes globally and across biomes. Most parameters have minimal impact on biophysical fluxes, with only a few substantially influencing results. While both models exhibit similar parameter sensitivity for some fluxes, CLM-FATES shows larger spread in gross primary productivity (GPP), driven by strong sensitivity to carboxylation rate. CLM-FATES also shows a weaker GPP response to soil hydrology parameters and exhibits higher water use efficiency (WUE). Cross-model comparisons reveal similar sensitivities for some parameters (e.g., leaf dimension) but divergent responses to others (e.g., stomatal intercept), highlighting underlying structural differences. Differences in WUE and sensitivity to hydrology and stomatal conductance parameters underscore how model structure fundamentally alters parametric sensitivity. The data sets generated from these ensembles can be used to identify influential parameters and guide future calibration efforts.

Foster, A. C. [NSF National Center for Atmospheric

Taming nuclear mass models with Gaussian processes

We propose a new set of nuclear mass predictions based on multiple theoretical mass models. By employing Gaussian process regression with the Matérn kernel, we achieved root-mean-square (rms) deviations below 100 keV for the training dataset. The best-performing mass models achieved rms deviations below 150 keV for the new precise mass data from AME2020, whereas the ensemble average showed robust performance across the nuclear chart. Our approach uniquely combines: (1) systematic refinement of eight mass models through their residuals, (2) physics-informed features, including magic numbers, nucleon parity numbers, neutron excess, and nuclear collectivity, and (3) theory-to-theory validation demonstrating robust extrapolation capability. We find that the Matérn kernel provides superior uncertainty quantification compared to the RBF kernel, with a length-scale analysis revealing enhanced inter-nuclei correlations. We provide complete mass predictions for all unknown nuclides in AME2020, offering valuable constraints for nuclear structure studies and astrophysical modeling when used with proper uncertainty propagation.

Gaussian processes

Ensemble Simulations on Leadership Computing Systems

Scientific productivity can be enhanced through workflow management tools, relieving large High Performance Computing (HPC) system users from the tedious tasks of scheduling and designing the complex computational execution of scientific applications. This paper presents a study on the usage of ensemble workflow tools to accelerate science using the Summit and Frontier supercomputing systems. The research aims to connect science domain simulations using Oak Ridge Leadership Computing Facility (OLCF) supercomputing platforms with ensemble workflow methods in order to accelerate HPC-enabled discovery and boost scientific impact. We present the coupling, porting and optimization of Radical-Cybertools on three applications: Chroma, NAMD and LAMMPS. The tools augment traditional HPC monolithic runs with a pilot scheduler. Lessons-learned are discussed for physics, biology and materials science applications. We discuss intrinsic limitations of coupling and porting ensemble workflow tools to applications that run on large HPC systems. The origins of technical challenges and their solutions developed during the implementation process are discussed. Data management strategies, OLCF’s policies for ensembles, and natively supported workflow tools are also summarized.

Georgiadou, Antigoni [ORNL] (ORCID:000000020977631

Spectral Data Fusion From Handheld Laser-Induced Breakdown Spectroscopy (LIBS) and X-ray Fluorescence (XRF) Analyzers for Improved Detection of Cerium in a Simulated Dispersal Accident

Here, this work implements a mid-level data fusion methodology on spectral data from handheld X-ray fluorescence and laser-induced breakdown spectroscopy analyzers to quantify plutonium surrogate (CeO 2 ) contamination in soil samples for the first time. Spectral data from each analyzer were used independently to train supervised machine learning regressions to predict Ce concentration. Fused features from both data sets were then used to train the same models, comparing prediction performance by evaluating model precision and sensitivity. Fusing principal component scores from the two sensors yielded an order of magnitude improvement in precision and sensitivity of predictions made with an artificial neural network, compared to predictions made by models trained on independent sensor data. As a result, a boosted ensemble trained on the fused spectral features yielded an ideal predictor with root-mean-squared error on the order of 10 –6 and calculated limit of detection order 10 –5 wt %.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

MINE: a new way to design genetics experiments for discovery

Abstract The Maximally Informative Next Experiment or MINE is a new experimental design approach for experiments, such as those in omics, in which the number of effects or parameters p greatly exceeds the number of samples n (p > n). Classical experimental design presumes n > p for inference about parameters and its application to p > n can lead to over-fitting. To overcome p > n, MINE is an ensemble method, which makes predictions about future experiments from an existing ensemble of models consistent with available data in order to select the most informative next experiment. Its advantages are in exploration of the data for new relationships with n < p and being able to integrate smaller and more tractable experiments to replace adaptively one large classic experiment as discoveries are made. Thus, using MINE is model-guided and adaptive over time in a large omics study. Here, MINE is illustrated in two distinct multiyear experiments, one involving genetic networks in Neurospora crassa and a second one involving a genome-wide association study in Sorghum bicolor as a comparison to classic experimental design in an agricultural setting.

Biochemistry & Molecular Biology

Modeling aerosol transmission spectra from n(λ) and k(λ) infrared optical constants measurements of organic liquids and solids

The effects of light scattering and refraction play significantly different roles for aerosols than for bulk materials, making it challenging to identify aerosolized chemicals using traditional spectral methods or spectral reference libraries. Due to a potentially infinite number of particle morphologies, sizes, and compositions, constructing a database of laboratory-measured aerosol spectra is not a practical solution. Here, as an alternative approach, the measured n / k optical vectors of two example organic materials (diethyl phthalate and D-mannitol) are used in combination with particle absorption / scattering theory (Mie theory and FDTD) and the Beer-Lambert law to generate a series of synthetic infrared transmission / scattered light spectra. The synthetic spectra show significant differences versus simple slab transmission spectra, even for small changes in particle size (e.g., 5 vs. 10 µm) for both single particles and ensembles, potentially serving as useful reference data for aerosol sensing. For spherical single particles with diameters of 1 to 10 µm, FDTD simulations predict changes in the magnitudes of spectral shifts and the shapes of the peaks vs. particle size with only small deviations from Mie theory predictions, yet reliably capture the direction of the shifts. Typical spectral peak shifts in the longwave infrared correspond to Δλ ∼0.20 µm (∼34 cm -1 ) when compared to corresponding slab transmission spectra. Additionally, synthetic spectra generated from the n / k values derived using two different methods (KBr pellet transmission and single-angle reflectance) are compared using the Mie theory model.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Quarterly Soil Core and Root Analyses from the Missouri Ozarks AmeriFlux (MOFLUX) Site, Ashland, Missouri, 2017-2023

This dataset contains quarterly soil core measurements from the Missouri Ozarks AmeriFlux (MOFLUX) site located at the University of Missouri’s Thomas H. Baskett Wildlife Research and Education Area near Ashland, Missouri. These data will be used to parameterize an ensemble of MOFLUX-optimized soil carbon-nitrogen models, used to simulate carbon (C) and nitrogen (N) cycling responses to future hydroclimatic scenarios and the trajectory of soil C stocks with concomitant forest decline. Beginning in 2017, eight soil cores were collected approximately quarterly near plot 1 of the southeast transect, near the automated soil respiration flux chambers, from 0–15 cm depth. Data are currently available through 2023 (2017-06-14 to 2023-11-13); additional observations will be appended to this dataset as they become available. Cores were analyzed for gravimetric moisture content, pH, total carbon and nitrogen, texture, microbial biomass carbon and nitrogen, and extractable dissolved organic carbon and nitrogen. This dataset contains one data file in comma separate (*.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (*.csv) format and a user guide in PDF (*.pdf) format.

54 ENVIRONMENTAL SCIENCES