Search NASASearch

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Unbinned extraction of $γ$ from $B\to DK$ with normalizing flows

We introduce an unbinned method for extracting the CKM angle $γ$ from the decay chain $B^\pm \to (D \to K_S π^+ π^-) K^\pm$ using normalizing flows (NFs). The NFs, trained on $D$ decay data, learn a faithful continuous representation of the amplitude and strong phase variation over the $D\to K_Sπ^+π^-$ Dalitz plot whose fidelity improves with increased data sample sizes. With this input, the $B$ decay data can be used to extract the parameters $r_B$, $δ_B$, and $γ$. We test the method on Monte Carlo generated data, where it successfully recovers the injected value of $γ$ within uncertainties. The present implementation propagates statistical uncertainties from finite training data via an ensemble of independently trained flows, and does not attempt to capture the effects of systematic experimental errors. We explore two versions of the method that differ in how the trigonometric constraint on phase variation is encoded, and comment on the possible extension to Bayesian NFs, which would provide direct uncertainty estimates on the learned densities without requiring ensemble training.

Grossman, Yuval [Cornell U., LEPP]

Downscaled Daily 1 km Climate Data (NEX-GDDP-CMIP6) for Southeast Texas. Full ensemble of downscaled CMIP6 climate projections at 1 km daily resolution.

For the SETx-UIFL, the daily NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP-CMIP6) dataset climate projections were downscaled from approximately 27 km to 1 km. The SETx dataset provides very high-resolution climate data for the historical period (1950–2014) and future scenarios derived from CMIP6 global models under the four Tier 1 Shared Socioeconomic Pathways (SSPs 1.26, 2.45, 3.70, and 5.85), developed for the IPCC Sixth Assessment Report. A subset of ten NEX-GDDP-CMIP6 models was selected to represent a balance of model families, climate sensitivities, and availability across scenarios, ensuring a diverse and reliable ensemble for regional analysis. Selected models: BCC-CSM2-MR, CESM2, CMCC-ESM2, CNRM-ESM2-1, EC-Earth3, FGOALS-g3, GFDL-CM4, MPI-ESM1-2-HR, MRI-ESM2-0, NorESM2-MM. Daily variables downscaled include tasmax, tasmin, tas, pr, hurs, huss, rsds, rlds, and sfcWind.

Persad, Geeta

Dataset for manuscript "Equipartition and the temperature of maximum density of TIP4P/2005 water"

We simulate TIP4P/2005 water in the temperature range of 257 K to 318 K with time-steps 0.25, 0.50, 1.00, 2.00, and 4.00 fs. The density-temperature behavior obtained using 0.25 or 0.50 fs are in excellent agreement with each other but differ from those obtained using time-steps that have been shown earlier to lead to a breakdown of equipartition. The temperature of maximum density (TMD) is 277.15 K with time-step 0.25 or 0.50 fs, but is shifted to progressively lower values for longer time-steps, a trend that holds for different thermostat/barostat combinations. Enhancing the water-water dispersion interaction, as has been recommended for simulating disordered proteins in TIP4P/2005, degrades the description of the liquid-vapor phase envelope. We present a simple physically transparent reasoning to highlight the separation of the time-scales between translational and rotational motion. We also develop a metric, Chi, that we term the equipartition anomaly, to detect equipartition violations in simulations that include molecules that are treated as rigid objects. Calculating Chi is shown to be straightforward and sensitive to equipartition violations. A key takeaway from this study is that using sufficiently short time-steps (less than or equal to 0.5 fs) to preserve equipartition is essential for obtaining meaningful liquid water properties and for producing reliable simulation data, as correct-ensemble sampling is fundamental to ensure reproducibility across codes and simulation alogrithms. The included dataset provides the raw data used in the preparation of the graphs noted in the manuscript.

36 MATERIALS SCIENCE

Exploring Water System Vulnerabilities in California's Central Valley Under the Late Renaissance Megadrought and Climate Change

Abstract California faces cycles of drought and flooding that are projected to intensify, but these extremes may impact water users across the state differently due to the region's natural hydroclimate variability and complex institutional framework governing water deliveries. To assess these risks, this study introduces a novel exploratory modeling framework informed by paleo and climate‐change based scenarios to better understand how impacts propagate through the Central Valley's complex water system. A stochastic weather generator, conditioned on tree‐ring data, produces a large ensemble of daily weather sequences conditioned on drought and flood conditions under the Late Renaissance Megadrought period (1550–1580 CE). Regional climate changes are applied to this weather data and drive hydrologic projections for the Sacramento, San Joaquin, and Tulare Basins. The resulting streamflow ensembles are used in an exploratory stress test using the California Food‐Energy‐Water System model, a highly resolved, daily model of water storage and conveyance throughout California's Central Valley. Results show that megadrought conditions lead to unprecedented reductions in inflows and storage at major California reservoirs. Both junior and senior water rights holders experience multi‐year periods of curtailed water deliveries and complete drawdowns of groundwater assets. When megadrought dynamics are combined with climate change, risks for unprecedented depletion of reservoir storage and sustained curtailment of water deliveries across multiple years increase. Asymmetries in risk emerge depending on water source, rights, and access to groundwater banks.

Gupta, Rohini S. [School of Civil and Environmenta

Stochastic Ensemble Generation for Improved Characterization of Representing Geologic Variability in a Reservoir: IBDP Case Study for SMART Initiative

This document is a poster covering the findings from activities on training data generation, specifically geologic ensemble generation. The generated geologic realizations captured the range of possible permeability distributions of the subsurface at the Illinois Basin - Decatur Project (IBDP) site, based on available well log variabilities. The percentages of reservoirs and baffles in the injection zone and a truncation of baffle permeability led to more variance in the simulations. This will be used to build forward modeling, history matching, and optimization workflows. The geologic realizations were also ranked according to dynamic measures of hydraulic diffusivity, and simulations confirm a greater contrast between the reservoir and the baffles during injection.

stochastic ensemble generation

Accelerating the Discovery of New, Single Phase High Entropy Ceramics via Active Learning

High-entropy ceramics have garnered interest due to their remarkable hardness, compressive strength, thermal stability, and fracture toughness; yet the discovery of new high-entropy ceramics (out of a tremendous number of possible elemental permutations) still largely requires costly, inefficient, trial-and-error experimental and computational approaches. The entropy forming ability (EFA) factor was recently proposed as a computational descriptor that positively correlates with the likelihood that a 5-metal high-entropy carbide (HECs) will form the desired single phase, homogeneous solid solution; however, discovery of new compositions is computationally expensive. If you consider 8 candidate metals, the HEC EFA approach uses 49 optimizations for each of the 56 unique 5-metal carbides, requiring a total of 2744 costly density functional theory calculations. Here, we describe an orders-of-magnitude more efficient active learning (AL) approach for identifying novel HECs. To begin, we compared numerous methods for generating composition-based feature vectors (e.g., magpie and mat2vec), deployed an ensemble of machine learning (ML) models to generate an average and distribution of predictions, and then utilized the distribution as an uncertainty. Here we then deployed an AL approach to extract new training data points where the ensemble of ML models predicted a high EFA value or was uncertain of the prediction. Our approach has the combined benefit of decreasing the amount of training data required to reach acceptable prediction qualities and biases the predictions toward identifying HECs with the desired high EFA values, which are tentatively correlated with the formation of single phase HECs. Using this approach, we increased the number of 5-metal carbides screened from 56 to 15,504, revealing 4 compositions with record-high EFA values that were previously unreported in the literature. Our AL framework is also generalizable and could be modified to rationally predict optimized candidate materials/combinations with a wide range of desired properties (e.g., mechanical stability, thermal conductivity).

36 MATERIALS SCIENCE

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES

ARM Trajectories Data Set Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility’s ARM Trajectories Data Set (ARMTRAJ) Value-Added Product (VAP) provides trajectory data sets initialized at ARM deployment coordinates and configured using ARM data sets. The four trajectory data sets support aerosol, cloud, and planetary boundary-layer research. Trajectory calculations use the Hybrid Single-Particle Lagrangian Integrated Trajectory (HYSPLIT) model informed by the European Centre for Medium-Range Weather Forecasts (ECMWF) fifth-generation atmospheric reanalysis (ERA5) data set at its highest spatial resolution (~31 km). HYSPLIT also runs at multiple initial starting locations surrounding ARM deployments (in latitude/longitude and/or vertical coordinates), facilitating an ensemble for each sample in the data sets. The ensemble mean and variability reported in ARMTRAJ improve the fidelity and provide uncertainty estimates of trajectory coordinates, thermodynamic properties, and other output fields.

54 ENVIRONMENTAL SCIENCES

ARM Trajectories Data Set Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility’s ARM Trajectories Data Set (ARMTRAJ) Value-Added Product (VAP) provides trajectory data sets initialized at ARM deployment coordinates and configured using ARM data sets. The six trajectory data sets support aerosol, cloud, planetary boundary layer, and related research (aerosol-cloud interactions, etc.), as well as studies using ARM Aerial Facility (AAF) and tethered balloon system (TBS) measurements. Trajectory calculations use the Hybrid Single-Particle Lagrangian Integrated Trajectory (HYSPLIT) model informed by the European Centre for Medium-Range Weather Forecasts (ECMWF) fifth-generation atmospheric reanalysis (ERA5) data set at its highest spatial resolution (~31 km). HYSPLIT also runs at multiple initial starting locations surrounding ARM deployments (in latitude/longitude and/or vertical coordinates), facilitating an ensemble for each sample in the data sets. The ensemble mean and variability reported in ARMTRAJ improve the fidelity and provide uncertainty estimates of trajectory coordinates, thermodynamic properties, and other output fields.

54 ENVIRONMENTAL SCIENCES

PNNL-ANL Hydrometeorological Super Ensemble

The current dataset contains data upload links to the following **hydrologic (water balance), river routing (water management), and hydropower simulation** data over CONUS: * Climate Forcing: **Livneh** (https://www.nature.com/articles/sdata201542) * Simulation Scenario: **Historical** * Simulation Period: **1971-2013 (1972-2013 for Hydropower)** * Simulation Models: **VIC (Variable Infiltration Capacity)**, **mosartwmpy (Model for Scale Adaptive River Transport-Water Management in Python)**, and **PNNL B1Hydro** * Output Format: **NetCDF** and **CSV**

Tidwell, Vincent C [Pacific Northwest National Lab

PNNL-ANL Hydrometeorological Super Ensemble

The current dataset contains data upload links to the following hydrologic (water balance), river routing (water management), and hydropower simulation data over CONUS: Climate Forcing: ClimRR (https://climrr.anl.gov/climrrdata) Simulation Scenario: Historical, Mid-Century, End-Century Simulation Period: 1995-2004, 2045-2054, 2085-2094 Simulation Models: VIC (Variable Infiltration Capacity), mosartwmpy (Model for Scale Adaptive River Transport-Water Management in Python), and PNNL B1Hydro Output Format: NetCDF and CSV

Tidwell, Vincent C [Pacific Northwest National Lab

Strong Lens Discoveries in DESI Legacy Imaging Surveys DR10 with Two Deep Learning Architectures

Abstract We have conducted a search for strong gravitational lensing systems in the Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Surveys Data Release 10 (DR10). This paper is the fourth in a series of searches. This is the first catalog of lens candidates covering nearly the entirety of the extragalactic sky south of declination δ ≈ +32 ∘ , all observed by DECam, covering ∼14,000 deg 2 . We impose a z -band magnitude cut of <20 in AB magnitude. We deploy a residual neural network and EfficientNet as an ensemble trained on a compilation of known lensing systems and high-grade candidates as well as nonlenses in the same footprint. The predictions from these two base models are aggregated using a meta-learner. After applying our ensemble to the survey data, we exclude known candidates and systems, and use our own visual inspection portal to rank images in the top 0.01 percentile of all neural network recommendations. We have found 811 lens candidates, five of which are confirmed through Euclid Quick Data Release (Q1). These include 484 new candidates in the Legacy Surveys DR9 footprint, all parts of which have been searched for strong lenses at least once before, either by our group or others. Combining the discoveries from this work with those from the first three papers in this series (335, 1210, and 1512), we have discovered a total of 3868 new candidates in the DESI Legacy Surveys.

Inchausti, Jose Carlos [University of San Francisc

Joint state-parameter estimation for the reduced fracture model via the united filter

Here, in this paper, we introduce an effective United Filter method for jointly estimating the solution state and physical parameters in flow and transport problems within fractured porous media. Fluid flow and transport in fractured porous media are critical in subsurface hydrology, geophysics, and reservoir geomechanics. Reduced fracture models, which represent fractures as lower-dimensional interfaces, enable efficient multi-scale simulations. However, reduced fracture models also face accuracy challenges due to modeling errors and uncertainties in physical parameters such as permeability and fracture geometry. To address these challenges, we propose a United Filter method, which integrates the Ensemble Score Filter (EnSF) for state estimation with the Direct Filter for parameter estimation. EnSF, based on a score-based diffusion model framework, produces ensemble representations of the state distribution without deep learning. Meanwhile, the Direct Filter, a recursive Bayesian inference method, estimates parameters directly from state observations. The United Filter combines these methods iteratively: EnSF estimates are used to refine parameter values, which are then fed back to improve state estimation. Numerical experiments demonstrate that the United Filter method surpasses the state-of-the-art Augmented Ensemble Kalman Filter, delivering more accurate state and parameter estimation for reduced fracture models. This framework also provides a robust and efficient solution for PDE-constrained inverse problems with uncertainties and sparse observations.

Bayesian inference

CO 2 storage site characterization using ensemble-based approaches with deep generative models

Estimating spatially distributed properties such as permeability from available sparse measurements is a great challenge in efficient subsurface CO 2 storage operations. In this paper, a deep generative model that can accurately capture complex subsurface structure is tested with an ensemble-based inversion method for accurate and accelerated characterization of CO 2 storage sites. We chose Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP) for its realistic reservoir property representation and Ensemble Smoother with Multiple Data Assimilation (ES-MDA) for its robust data fitting and uncertainty quantification capability. WGAN-GP are trained to generate high-dimensional permeability fields from a low-dimensional latent space and ES-MDA then updates the latent variables by assimilating available measurements. Several subsurface site characterization examples including Gaussian, channelized, and fractured reservoirs are used to evaluate the accuracy and computational efficiency of the proposed method and the main features of the unknown permeability fields are characterized accurately with reliable uncertainty quantification. Furthermore, the estimation performance is compared with a widely-used variational, i.e., optimization-based, inversion approach, and the proposed approach outperforms the variational inversion method in several benchmark cases. We explain such superior performance by visualizing the objective function in the latent space: because of nonlinear and aggressive dimension reduction via generative modeling, the objective function surface becomes extremely complex while the ensemble approximation can smooth out the multi-modal surface during the minimization. This suggests that the ensemble-based approach works well over the variational approach when combined with deep generative models at the cost of forward model runs unless convergence-ensuring modifications are implemented in the variational inversion.

42 ENGINEERING

A Greening Future Elevates Flash Drought Risk in Northern Mid‐to‐High Latitudes

Flash droughts have become a growing concern, as they can emerge rapidly and increase the risk of crop failure. Although past studies have investigated the meteorological drivers and future changes of flash drought, why flash drought is more frequent over humid and vegetated regions remains underexplored. This study delves further into the mechanism by which vegetation regulates flash drought and its future change using observations from multiple data sets and large ensemble simulations from three Earth system models. On an interannual timescale, both observations and simulations show robust increases in flash drought frequency and a higher flash-to-sub-seasonal drought ratio during spring or antecedent conditions with dense vegetation, supporting the important role of vegetation in flash drought occurrence, especially in the northern mid-to-high latitudes. In the latter regions, the large ensemble simulations show robust increases in flash drought (e.g., 67% and 46% increases in Eastern U.S. and North Asia in 2050–2100 relative to 1950–2000 under the high emission scenario), where the growing season is lengthening. Although greening might suggest reduced drought stress, it drives precipitation-soil moisture-evapotranspiration decoupling by increasing evapotranspiration partitioning to transpiration. As transpiration can access deep soil water through the plant root system, its increased portion can weaken the constraints of concurrent precipitation on evapotranspiration, thus accelerating soil moisture depletion under high evaporative demand, driving a slow-to-rapid drought transition. How vegetation regulates flash drought by regulating surface moisture budget is supported by observations and simulations. Although warming supports early planting, agriculture may increasingly be threatened by surging flash drought risk.

Drought

How Well Can CMIP6 Models Represent the Observed Influence of the Pacific and Indian Oceans on the Indian Summer Monsoon Rainfall?

This study evaluates the ability of CMIP6 climate models to simulate the observed effects of tropical Pacific and Indian Ocean sea surface temperature anomalies (SSTAs) on Indian summer monsoon rainfall (ISMR) variability. Using observational data and the large ensemble historical simulations of seven CMIP6 models from 1950 to 2014, we applied a cyclostationary linear inverse model (CS-LIM) to isolate the impacts of tropical Pacific SSTAs, Indian Ocean SSTAs and their interaction on the interannual variability of ISMR. Overall, CMIP6 models well reproduced the observed enhanced (reduced) ISMR variability from Pacific SSTAs (Indian Ocean SSTAs and the Indo-Pacific interaction), but with varying spatial patterns and magnitudes. While CESM2 and E3SM-2-0 showed the best agreement with observations for the effects of Pacific SSTAs and the Indo-Pacific interaction, respectively, CMIP6 models showed mixed results for the impacts from Indian Ocean SSTAs. Composite analysis of ISMR anomalies during the developing phases of pure and co-occurring El Niño-Southern Oscillation (ENSO) and Indian Ocean dipole (IOD) events revealed that the impacts from Pacific SSTAs were captured reasonably well by E3SM-2-0, CESM2, MIROC6, and MPI-ESM1-2-LR, while E3SM-2-0 also showed the best agreement with observations for the effects from the Indo-Pacific interaction. However, all models showed substantial biases in simulating the Indian Ocean SSTA impacts on ISMR, especially for pure El Niño events. Overall, this study provides new insights into how individual CMIP6 models simulate the isolated impacts from the tropical Pacific and Indian Oceans, which has important applications for improving ISMR predictions and interpreting ISMR future projections.

monsoon

Wilkins: HPC in situ workflows made easy

In situ approaches can accelerate the pace of scientific discoveries by allowing scientists to perform data analysis at simulation time. Current in situ workflow systems, however, face challenges in handling the growing complexity and diverse computational requirements of scientific tasks. In this work, we present Wilkins, an in situ workflow system that is designed for ease-of-use while providing scalable and efficient execution of workflow tasks. Wilkins provides a flexible workflow description interface, employs a high-performance data transport layer based on HDF5, and supports tasks with disparate data rates by providing a flow control mechanism. Wilkins seamlessly couples scientific tasks that already use HDF5, without requiring task code modifications. We demonstrate the above features using both synthetic benchmarks and two science use cases in materials science and cosmology.

HPC

LQCD NN S-wave @ SU(3) on C103

This data contains lattice QCD data files for one- and two-nucleon correlation functions on the C103 ensemble. It also contains lattice QCD data sets needed to compute the HAL QCD potential on the same configurations. 'LQCD NN S-wave @ SU(3) on C103' Copyright (c) 2025, The Regents of the University of California, through Lawrence Berkeley National Laboratory (subject to receipt of any required approvals from the U.S. Dept. of Energy). All rights reserved. This work is openly licensed via CC BY 4.0.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS