Search NASASearch

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

On learning what to learn: Heterogeneous observations of dynamics and establishing possibly causal relations among them

Abstract Before we attempt to (approximately) learn a function between two sets of observables of a physical process, we must first decide what the inputs and outputs of the desired function are going to be. Here we demonstrate two distinct, data-driven ways of first deciding “the right quantities” to relate through such a function, and then proceeding to learn it. This is accomplished by first processing simultaneous heterogeneous data streams (ensembles of time series) from observations of a physical system: records of multiple observation processes of the system. We determine (i) what subsets of observables are common between the observation processes (and therefore observable from each other, relatable through a function); and (ii) what information is unrelated to these common observables, therefore particular to each observation process, and not contributing to the desired function. Any data-driven technique can subsequently be used to learn the input–output relation—from k-nearest neighbors and Geometric Harmonics to Gaussian Processes and Neural Networks. Two particular “twists” of the approach are discussed. The first has to do with the identifiability of particular quantities of interest from the measurements. We now construct mappings from a single set of observations from one process to entire level sets of measurements of the second process, consistent with this single set. The second attempts to relate our framework to a form of causality: if one of the observation processes measures “now,” while the second observation process measures “in the future,” the function to be learned among what is common across observation processes constitutes a dynamical model for the system evolution.

Sroczynski, David W.

Using Multiple Isotope-Labeled Infrared Spectra for the Structural Characterization of an Intrinsically Disordered Peptide

Intrinsically disordered proteins (IDPs) rapidly interconvert between conformers, requiring an ensemble description. This complicates their experimental characterization, and force field limitations pose challenges for their simulation. Here, in this work, we use isotope-labeled and unlabeled infrared (IR) spectra to reweight simulated ensembles of the elastin-like peptide GVGVPGVG, a paradigmatic disordered peptide. By comparing the results obtained with different spectra, we explicitly show that the weights are underdetermined by the ensemble averaged data. We identify which labels and frequency regions maximize structural information while minimizing sensitivity to simulation error and show that these regions report on whether the peptide makes specific interactions. Our work shows the importance of incorporating simulations and simulated spectra at the planning stages of isotope-labeled IR experiments and more generally provides a framework for interpreting IR data for IDPs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Dark Matter Velocity Distributions for Direct Detection: Astrophysical Uncertainties Are Smaller Than They Appear

The sensitivity of direct detection experiments depends on the phase-space distribution of dark matter near the Sun, which can be modeled theoretically using cosmological hydrodynamical simulations of Milky Way–like galaxies. However, capturing the halo-to-halo variation in the local dark matter speeds—a necessary step for quantifying the astrophysical uncertainties that feed into experimental results—requires a sufficiently large sample of simulated galaxies, which has been a challenge. In this Letter, we quantify this variation with nearly 100 Milky Way–like galaxies from the tng50 simulation, the largest sample to date at this resolution. Moreover, we introduce a novel phase-space scaling procedure that endows every system with a reference frame that accurately reproduces the local standard-of-rest speed of our Galaxy, providing a principled way of extrapolating the simulation results to real-world data. The ensemble of predicted speed distributions is well characterized by the standard halo model, a Maxwell-Boltzmann distribution truncated at the escape speed, though the individual distributions can deviate from it, especially at high speeds. The dark matter–nucleon cross section limits placed by these speed distributions vary by ∼ 60% about the median. This places the 1⁢𝜎 astrophysical uncertainty at or below the level of the systematic uncertainty of current ton-scale detectors, even down to the energy threshold. The predicted uncertainty remains unchanged when subselecting on those TNG 50 galaxies with merger histories similar to the Milky Way. Tabulated speed distributions, as well as Maxwell-Boltzmann fits, are provided for use in computing direct detection bounds or projecting sensitivities.

Milky Way

Changing effects of external forcing on Atlantic–Pacific interactions

Recent studies have highlighted the increasingly dominant role of external forcing in driving Atlantic and Pacific Ocean variability during the second half of the 20th century. This paper provides insights into the underlying mechanisms driving interactions between modes of variability over the two basins. We define a set of possible drivers of these interactions and apply causal discovery to reanalysis data, two ensembles of pacemaker simulations where sea surface temperatures in either the tropical Pacific or the North Atlantic are nudged to observations, and a pre-industrial control run. We also utilize large-ensemble means of historical simulations from the Coupled Model Intercomparison Project Phase 6 (CMIP6) to quantify the effect of external forcing and improve the understanding of its impact. A causal analysis of the historical time series between 1950 and 2014 identifies a regime switch in the interactions between major modes of Atlantic and Pacific climate variability in both reanalysis and pacemaker simulations. A sliding window causal analysis reveals a decaying El Niño–Southern Oscillation (ENSO) effect on the Atlantic as the North Atlantic fluctuates towards an anomalously warm state. The causal networks also demonstrate that external forcing contributed to strengthening the Atlantic's negative-sign effect on ENSO since the mid-1980s, where warming tropical Atlantic sea surface temperatures induce a La Niña-like cooling in the equatorial Pacific during the following season through an intensification of the Pacific Walker circulation. The strengthening of this effect is not detected when the historical external forcing signal is removed in the Pacific pacemaker ensemble. The analysis of the pre-industrial control run supports the notion that the Atlantic and Pacific modes of natural climate variability exert contrasting impacts on each other even in the absence of anthropogenic forcing. The interactions are shown to be modulated by the (multi)decadal states of temperature anomalies of both basins with stronger connections when these states are “out of phase”. We show that causal discovery can detect previously documented connections and provides important potential for a deeper understanding of the mechanisms driving changes in regional and global climate variability.

54 ENVIRONMENTAL SCIENCES

MINE: maximally informative next experiment—toward a new GWAS experimental design and methodology

Abstract The computational methodology of Genome Wide Association Studies (GWAS) currently has several limitations: (i) the number of observations (rows) on a quantitative trait tends to be smaller than the number of single nucleotide polymorphisms (SNPs) (columns) in the design matrix; (ii) each SNP is usually modeled separately, failing to acknowledge interaction between each other (ie epistasis); (iii) there is implicit linkage disequilibrium (LD) between neighboring SNPs due to their linkage. To overcome these issues, we developed a tool that uses ensemble methods to fit mixed linear models to GWAS data, and these ensemble methods include the development of a new experimental design approach in GWAS, which uses the resultant models and data to select the next informative experiment over time. This new adaptive and staged approach for GWAS experimental design was developed and tested in a 3 yr adaptive model-guided discovery experiment against a fixed classical design. In Sorghum bicolor a total of 79, 86, and 78 accessions were tested in years 1, 2, and 3, respectively out of 343 accessions available in the Bioenergy Association Panel (BAP) each identified for 232,303 SNPs, 1 every 2–3 kb in the genomes. We demonstrated the feasibility of MINE enacted with 8 people in the field per year over 3 yr vs in 1 large classical design enacted with 20 people in 1 yr. The MINE results for chromosomal regions identified controlling dry weight were confirmed against results from previous sorghum GWAS experiments and 1 large classical design for the BAP panel.

Genetics & Heredity

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS

Simultaneous inference of equation of state parameters and unknown data errors with uncertainty quantification via hierarchical Bayesian posterior maximization

Equations of state (EOSs) are a key component in running hydrodynamic simulations as they relate the thermodynamic states for the material. The Davis reactants EOS is commonly used for modeling high explosives (HEs), and the EOS model parameters are calibrated using material specific data. The calibrations are often performed with uncertainty quantification via Bayesian inference to account for uncertainty in the data and generate ensembles of likely parameters. However, there are relatively few HE data sets to use for calibration and many are historical and lack error information. In this work, we simultaneously calibrate the Davis reactants EOS model parameters and unknown data error terms for the high explosive PBX 9501. To quantify the uncertainty in the models and the data, we use a Bayesian framework for the calibration and compute the hierarchical Bayesian posterior distribution with both a posteriori maximization approach and Markov Chain Monte Carlo. In general, we find that, given our assumptions, the two approaches result in similar calibrated parameters, posterior covariance matrices, and insights about the parameters but that the posterior maximization requires far less computational resources.

97 MATHEMATICS AND COMPUTING

Event Classifications on DNE2 Main Experiment Data using a Convolutional Neural Network Ensemble

The Dynamic Networks (DN) Experiment for FY24 (DNE2) is an experiment within DN with the goal of quantitatively evaluating the effectiveness of solutions developed so far by various researchers under the Low Yield Nuclear Monitoring (LYNM) program using a shared set of metrics and datasets. A key component of this experiment is the mimicking of a signature processing pipeline, and comparing currently accepted and standard-use processing methods to more state-of-the-art processes developed under DN. In this work, we focus specifically on the Event Characterization (EC) Focus Area (FA) of the pipeline, where a seismic event’s magnitude, yield and class are identified. We use Deep Learning (DL) to classify the type of events being processed as either earthquakes (EQs) or explosions (EXs) for three iterations of experiment datasets. The model is noticeably more confident and accurate in classifying explosions than earthquakes, reflecting a known shortcoming of the model, that being of a bias towards predicting explosions over earthquakes in the west coast due to training data biases.

97 MATHEMATICS AND COMPUTING

Boosting efficiency and reducing graph reliance: Basis adaptation integration in Bayesian multi-fidelity networks

The computational cost of high-fidelity numerical models makes outer-loop analysis, which requires repeated interrogation of the model such as uncertainty quantification, computationally demanding. Multi-fidelity methods, which construct a surrogate model using data from an ensemble of models of varying cost and accuracy, can substantially reduce the cost of outer-loop analysis. However, these methods can be difficult to apply when the model ensemble does not admit a clear hierarchy a priori and the correlations between models are low. Consequently, in this paper, we present a multi-fidelity method that leverages dimension reduction to enhance the correlation between models, thereby reducing the amount of data needed to train a surrogate from an unordered ensemble of models. Our method utilizes basis adaptation to build low-dimensional polynomial chaos expansions of each model and employs Multi-fidelity Networks to encode the relationships among models. We show that the resulting method exhibit two notable advantages over its counterpart: (1) enhanced accuracy (both reduced bias and variance); and (2) reduced dependency on the graph structure encoding relationships among models. We demonstrate the approach on an analytical test problem and a challenging finite element model for a spent nuclear fuel. Our method produces a surrogate model that is significantly more accurate than either a single-fidelity surrogate or a multi-fidelity surrogate constructed without basis adaptation.

42 ENGINEERING

Ensemble variational Fokker-Planck methods for data assimilation

Particle flow filters solve Bayesian inference problems by smoothly transforming a set of particles into samples from the posterior distribution. Particles move in state space under the flow of an McKean-Vlasov-Itˆo process. This work introduces the Variational Fokker-Planck (VFP) framework for data assimilation, a general approach that includes previously known particle flow filters as special cases. The McKean-Vlasov-Itˆo process that transforms particles is defined via an optimal drift that depends on the selected diffusion term. It is established that the underlying probability density - sampled by the ensemble of particles - converges to the Bayesian posterior probability density. For a finite number of particles the optimal drift contains a regularization term that nudges particles toward becoming independent random variables. Based on this analysis, we derive computationally-feasible approximate regularization approaches that penalize the mutual information between pairs of particles, and avoid particle collapse. Moreover, the diffusion plays a role akin to a particle rejuvenation approach that aims to alleviate particle collapse. The VFP framework is very flexible. Different assumptions on prior and intermediate probability distributions can be used to implement the optimal drift, and localization and covariance shrinkage can be applied to alleviate the curse of dimensionality. A robust implicit-explicit method is discussed for the efficient integration of stiff McKean- Vlasov-Itˆo processes. Here, the effectiveness of the VFP framework is demonstrated on three progressively more challenging test problems, namely the Lorenz ’63, Lorenz ’96 and the quasi-geostrophic equations.

97 MATHEMATICS AND COMPUTING

Convection-Permitting Ensembles of an Isolated Mountain Thunderstorm during RELAMPAGO/CACTI

Abstract The north–south-oriented Sierras de Córdoba (SDC) ridge in central Argentina is noted for initiating thunderstorms that may grow into intense mesoscale convective systems (MCSs). It also initiates more isolated, shorter-lived cells under weaker synoptic forcing. These cells are less impactful than MCSs but may be difficult to predict in convective-scale numerical weather prediction (NWP) due to their strong sensitivities to subgrid and partially resolved processes. To study the mechanisms and predictability of such cells, convection-permitting ensemble simulations were conducted of an isolated, diurnally forced SDC thunderstorm during Cloud, Aerosol, and Complex Terrain Interactions (CACTI)/Remote Sensing of Electrification, Lightning, and Mesoscale/Microscale Processes with Adaptive Ground Observations (RELAMPAGO). The rich observational data facilitated detailed ensemble verification, where dry biases in the surface energy balance and soil moisture were identified. These biases promoted rapid removal of convective inhibition and an early onset of precipitating cells over the SDC that were shallower and weaker than the observed cell. Correction, and then overcorrection, of the soil moisture bias in two successive ensembles was required to rectify the surface energy balance and improve the representation of the SDC cell. Nevertheless, substantial ensemble variability in convective precipitation was found, with some members producing more widespread convection than observed and others producing no deep convection at all. This variability was largely explained by a combination of thermodynamic and dynamic mechanisms, dominated by a positive sensitivity of convective precipitation to preconvective moist instability over the ridge. Secondary sensitivities were found to low-level upward mass flux and midlevel cross-barrier winds, the latter of which caused gravity waves with elevated downdrafts that tended to suppress incipient clouds.

Lopez, Andres [Department of Atmospheric and Ocean

Data from: Coupled machine learning-ecosystem ensemble models substantially improve predictions of nitrous oxide (N 2 O) fluxes from US croplands

Nitrous oxide (N₂O) is a potent and persistent greenhouse gas, with rising atmospheric concentrations driven in part by inefficient use of synthetic nitrogen (N) fertilizers in agriculture. Predicting soil N₂O emissions is challenging due to high spatial and temporal variability arising from complex soil biogeochemical processes. Process-based ecosystem models and standalone machine learning (ML) approaches without extensive site-specific calibration often miss high emission episodes. Here, we show how an Ensemble Modeling System (EMS) based on outputs from an ensemble of ecosystem models coupled to an ensemble of ML models can improve predictions and understanding of N2O fluxes from US cropland. Trained and validated on approximately 12,000 N2O chamber measurements at 17 U.S. Midwest sites (six crops, 35 management practices), the EMS accurately predicted daily fluxes of N2O at both training (R² = 0.84, RMSE = 16.4 g N ha⁻¹ d⁻¹) and held-out testing sites (R² = 0.84, RMSE = 6.2 g N ha⁻¹ d⁻¹). Analyses identified six dominant N₂O drivers: soil organic carbon (SOC), NH₄⁺, NO₃⁻, water-filled pore space (WFPS), soil temperature, and biomass production. Wet, warm soils produced large N₂O peaks only with sufficient SOC and mineral N; in low-SOC soils, fluxes remained low. Incorporating these drivers into process-based models might significantly improve their predictive capacity. The EMS demonstrates a strong potential to predict N₂O fluxes at unseen sites, enabling more reliable regional inventories, improved gap-filling where measurements are sparse, and enhanced understanding of mechanisms to advance targeted mitigation strategies in food, feed, and bioenergy crops.

agricultural sciences

Implementation of stacked ensemble machine learning for the detection of surrogate plutonium contamination in soil via LIBS

Supervised machine learning methods have demonstrated increased utility for the quantification of lanthanide and actinide elements in atomic spectroscopy applications. This study implements laser-induced breakdown spectroscopy (LIBS) for the identification of plutonium surrogate material (CeO 2 ) in soil matrices by training supervised machine learning methods on the recorded spectral data. A bagged ensemble using Random Forest yields the highest sensitivity predictions with a detection limit of 0.015 wt.% CeO 2 . However, high precision in Ce content prediction required the use of a stacked ensemble regression, which provided the superlative Ce quantification model with an error of 0.107% and a detection limit of 0.022 wt.%. Furthermore, the high performance of the stacked ensemble demonstrates its potential to enhance the accuracy and sensitivity of nuclear contaminant detection using field-deployable spectroscopic analyzers in real-world scenarios.

47 OTHER INSTRUMENTATION

Equipartition and the Temperature of Maximum Density of TIP4P/2005 Water

Here, we simulate TIP4P/2005 water in the temperature range of 257 to 318 K with time-steps δ = 0.25, 0.50, 1.00, 2.00, and 4.00 fs. The density–temperature behavior obtained using 0.25 or 0.50 fs is in excellent agreement with each other but differs from those obtained using time steps that have been shown earlier to lead to a breakdown of equipartition. For δt = 0.25 or 0.50 fs, the temperature of maximum density (TMD) is 277.15 K and the density value is in close agreement with experiments. For δt = 1.00 fs, the TMD is 277.15 K, but the density value is shifted higher. For the other time steps considered here, the TMD is shifted to progressively lower values for longer time steps, a trend that holds for different thermostat/barostat combinations. Enhancing the water–water dispersion interaction, as has been recommended for simulating disordered proteins in TIP4P/2005, degrades the description of the liquid–vapor phase envelope. We present a simple physically transparent explanation that highlights the separation of the time scales between translational and rotational motion. We also develop a metric, χ, that we term the equipartition anomaly, to detect equipartition violations in simulations that include molecules that are treated as rigid objects. Calculating χ is shown to be straightforward and sensitive to equipartition violations. A key takeaway from this study is that using sufficiently short time steps (≤0.5 fs) to preserve equipartition is essential for obtaining meaningful liquid water properties and for producing reliable simulation data, as correct ensemble sampling is fundamental to ensure reproducibility across codes and simulation algorithms.

Asthagiri, Dilipkumar N. [Oak Ridge National Labor

Unbinned extraction of $γ$ from $B\to DK$ with normalizing flows

We introduce an unbinned method for extracting the CKM angle $γ$ from the decay chain $B^\pm \to (D \to K_S π^+ π^-) K^\pm$ using normalizing flows (NFs). The NFs, trained on $D$ decay data, learn a faithful continuous representation of the amplitude and strong phase variation over the $D\to K_Sπ^+π^-$ Dalitz plot whose fidelity improves with increased data sample sizes. With this input, the $B$ decay data can be used to extract the parameters $r_B$, $δ_B$, and $γ$. We test the method on Monte Carlo generated data, where it successfully recovers the injected value of $γ$ within uncertainties. The present implementation propagates statistical uncertainties from finite training data via an ensemble of independently trained flows, and does not attempt to capture the effects of systematic experimental errors. We explore two versions of the method that differ in how the trigonometric constraint on phase variation is encoded, and comment on the possible extension to Bayesian NFs, which would provide direct uncertainty estimates on the learned densities without requiring ensemble training.

Grossman, Yuval [Cornell U., LEPP]

Downscaled Daily 1 km Climate Data (NEX-GDDP-CMIP6) for Southeast Texas. Full ensemble of downscaled CMIP6 climate projections at 1 km daily resolution.

For the SETx-UIFL, the daily NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP-CMIP6) dataset climate projections were downscaled from approximately 27 km to 1 km. The SETx dataset provides very high-resolution climate data for the historical period (1950–2014) and future scenarios derived from CMIP6 global models under the four Tier 1 Shared Socioeconomic Pathways (SSPs 1.26, 2.45, 3.70, and 5.85), developed for the IPCC Sixth Assessment Report. A subset of ten NEX-GDDP-CMIP6 models was selected to represent a balance of model families, climate sensitivities, and availability across scenarios, ensuring a diverse and reliable ensemble for regional analysis. Selected models: BCC-CSM2-MR, CESM2, CMCC-ESM2, CNRM-ESM2-1, EC-Earth3, FGOALS-g3, GFDL-CM4, MPI-ESM1-2-HR, MRI-ESM2-0, NorESM2-MM. Daily variables downscaled include tasmax, tasmin, tas, pr, hurs, huss, rsds, rlds, and sfcWind.

Persad, Geeta