Search NASA⌕ Search

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns

Coupled AI-Simulation workflows are becoming the major workloads for HPC facilities, and their increasing complexity necessitates new tools for performance analysis and prototyping of new in-situ workflows. We present SimAI-Bench, a tool designed to both prototype and evaluate these coupled workflows. In this paper, we use SimAI-Bench to benchmark the data transport performance of two common patterns on the Aurora supercomputer: a one-to-one workflow with co-located simulation and AI training instances, and a many-to-one workflow where a single AI model is trained from an ensemble of simulations. For the one-to-one pattern, our analysis shows that node-local and DragonHPC data staging strategies provide excellent performance compared Redis and Lustre file system. For the many-to-one pattern, we find that data transport becomes a dominant bottleneck as the ensemble size grows. Our evaluation reveals that file system is the optimal solution among the tested strategies for the many-to-one pattern.

Tummalapalli, Harikrishna [Argonne National Labora↗

Precipitation and Carbon-Water Coupling Jointly Control the Interannual Variability of Global Land Gross Primary Production

Carbon uptake by terrestrial ecosystems is increasing along with the rising of atmospheric CO2 concentration. Embedded in this trend, recent studies suggested that the interannual variability (IAV) of global carbon fluxes may be dominated by semi-arid ecosystems, but the underlying mechanisms of this high variability in these specific regions are not well known. Here we derive an ensemble of gross primary production (GPP) estimates using the average of three data-driven models and eleven process-based models. These models are weighted by their spatial representativeness of the satellite-based solar-induced chlorophyll fluorescence (SIF). We then use this weighted GPP ensemble to investigate the GPP variability for different aridity regimes. We show that semi-arid regions contribute to 57% of the detrended IAV of global GPP. Moreover, in regions with higher GPP variability, GPP fluctuations are mostly controlled by precipitation and strongly coupled with evapotranspiration (ET). This higher GPP IAV in semi-arid regions is co-limited by supply (precipitation)-induced ET variability and GPP-ET coupling strength. Our results demonstrate the importance of semi-arid regions to the global terrestrial carbon cycle and posit that there will be larger GPP and ET variations in the future with changes in precipitation patterns and dryland expansion.

chlorophyll fluorescence↗

Machine learning materials properties with accurate predictions, uncertainty estimates, domain guidance, and persistent online accessibility

One compelling vision of the future of materials discovery and design involves the use of machine learning (ML) models to predict materials properties and then rapidly find materials tailored for specific applications. However, realizing this vision requires both providing detailed uncertainty quantification (model prediction errors and domain of applicability) and making models readily usable. At present, it is common practice in the community to assess ML model performance only in terms of prediction accuracy (e.g. mean absolute error), while neglecting detailed uncertainty quantification and robust model accessibility and usability. Here, we demonstrate a practical method for realizing both uncertainty and accessibility features with a large set of models. We develop random forest ML models for 33 materials properties spanning an array of data sources (computational and experimental) and property types (electrical, mechanical, thermodynamic, etc). All models have calibrated ensemble error bars to quantify prediction uncertainty and domain of applicability guidance enabled by kernel-density-estimate-based feature distance measures. All data and models are publicly hosted on the Garden-AI infrastructure, which provides an easy-to-use, persistent interface for model dissemination that permits models to be invoked with only a few lines of Python code. We demonstrate the power of this approach by using our models to conduct a fully ML-based materials discovery exercise to search for new stable, highly active perovskite oxide catalyst materials.

domain of applicability↗

Dissemination of Global Surface Water Mapping from SAR and Optical Data to Global Stakeholders

Under NASA Solicitation A.37 “Earth Science Applications: Disaster Risk Reduction and Response,” NASA’s Disasters Program is funding projects focused on flood forecasting, post-event flood mapping, flood depth estimation and pre-event flood severity using Earth observation (EO) datasets (SAR and optical imagery) and derived flood products. A new initiative in the Disasters Program is underway to disseminate the flood products from different sensors to global stakeholders via the Pacific Disaster Center’s DisasterAWARE®, the NASA Disasters Mapping Portal and potentially other mechanisms. This initiative will also generate an integrated flood product(s) using an ensemble approach that will combine outputs from hydrologic models and EO data for different stakeholders globally for decision-making and response efforts.

remote sensing↗

Cataclysmic variable evolution - Clues from the underlying white dwarf

This paper presents an update of determinations of the CV white dwarf effective-temperature, T(eff), together with an initial exploration of the possible implications and constraints on the CV lifetimes and evolution based on the ensemble of white dwarf T(eff) values as a function of orbital period. The CV dwarf luminosities are derived by using the T(eff) data and adopting the masses of individual CV white dwarfs determined by Webbink (1990). The present ensemble of empirically determined white dwarf effective temperatures reveals a distribution centered near 16,000 K, implying a mean lower limit total cooling lifetime of 5 x 10 to the 8th yr for the majority of CV degenerates. The two coolest CV degenerates, VV Puppis and St LMi, were found among the strongly magnetic AM Her CVs.

Sion, Edward M.↗

Parallel Consensual Neural Networks

A new neural network architecture is proposed and applied in classification of remote sensing/geographic data from multiple sources. The new architecture is called the parallel consensual neural network and its relation to hierarchical and ensemble neural networks is discussed. The parallel consensual neural network architecture is based on statistical consensus theory. The input data are transformed several times and the different transformed data are applied as if they were independent inputs and are classified using stage neural networks. Finally, the outputs from the stage networks are then weighted and combined to make a decision. Experimental results based on remote sensing data and geographic data are given. The performance of the consensual neural network architecture is compared to that of a two-layer (one hidden layer) conjugate-gradient backpropagation neural network. The results with the proposed neural network architecture compare favorably in terms of classification accuracy to the backpropagation method.

Benediktsson, J. A.↗

Openet: Applications of Satellite-Based Evapotranspiration Data for Water Resources Management in the Western United States

Advancing water security in overallocated river basins globally requires consistent and reproducible information on consumptive use of water that can anchor the development of data-driven solutions to the challenge of balancing water supply and demand. OpenET is a fully automated system for field-scale (30 m), satellite-based mapping of evapotranspiration (ET) at daily, monthly and annual timesteps. OpenET currently provides spatially contiguous data throughout the 23 westernmost states in the continental US, and includes both current information as well as multi-year timeseries of ET. The OpenET consortium has implemented an ensemble of satellite-based ET models (ALEXI/DisALEXI, eeMETRIC, PT-JPL, geeSEBAL, SIMS and SSEBop) on Google Earth Engine, which provides a shared computing platform for collaboration on processing of data from Landsat and other satellites, land cover and meteorological inputs, leading to increased consistency and accuracy across the ensemble of models. Earth Engine also facilitates hosting and distribution of data via open data collections and an application programming interface. We provide updates on the OpenET framework, open data services and data access tools, approach to geographic expansion, recent accuracy assessments, and describe how a user-driven design approach has facilitated successful applications of OpenET data for a wide range of water resource management activities. Applications to date include: use of ET data to improve quantification of ET and consumptive use in Oregon, Utah and the Upper Colorado River Basin; streamlining of water use reporting requirements in the California Delta; support for calculation of water budgets for the implementation of the Sustainable Groundwater Management Act in California; and integration into decision support tools for irrigation management. The use cases demonstrate how satellite-derived ET data that are easily accessed and seen as broadly accepted can accelerate adoption of innovative water management practices at scale, and support advances in the sustainability of water supplies. Uptake and use of data by the OpenET science community has also led to advances in our understanding of the impacts of landcover change, irrigation intensification and wildfire events on hydrology and the water security.

Applications↗

Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure

First-principles fusion plasma simulations are both compute and memory intensive, and CGYRO is no exception. The use of many HPC nodes to fit the problem in the available memory thus results in significant communication overhead, which is hard to avoid for any single simulation. That said, most fusion studies are composed of ensembles of simulations, so we developed a new tool, named XGYRO, that executes a whole ensemble of CGYRO simulations as a single HPC job. By treating the ensemble as a unit, XGYRO can alter the global buffer distribution logic and apply optimizations that are not feasible on any single simulation, but only on the ensemble as a whole. The main saving comes from the sharing of the collisional constant tensor structure, since its values are typically identical between parameter-sweep simulations. This data structure dominates the memory consumption of CGYRO simulations, so distributing it among the whole ensemble results in drastic memory savings for each simulation, which in turn results in overall lower communication overhead.

CGYRO↗

Monitoring Extreme Weather in the Hindu Kush Himalaya Region

Why is monitoring extreme weather events important? The HKH (Hindu Kush Himalaya region experiences many extreme weather events, such as thunderstorms, especially during monsoon season. These events can cause economic hardship and loss of life. Monitoring Extreme Weather in the HKH Region is a service in development through SERVIR-Hindu Kush Himalaya that aims to develop a customized numerical weather prediction toolkit to assess these high impact events in this relatively data-sparse region. The High Impact Weather Assessment Toolkit (HIWAT) consists of an ensemble Weather Research and Forecasting (WRF)model, threat assessments based on the Global Precipitation Measurement (GPM) missions, and impact assessments based on Landsat and the Moderate Resolution Imaging Spectroradiometer (MODIS) imagery. In spring 2019, we began validation of forecasted precipitation using station data in Bangladesh and Climate Hazards Group InfraRed with Station data (CHIRPS).

Remote Sensing↗

Non-stationary precipitation design standards for stormwater infrastructure modernization at USAF installations

The resilience of defense infrastructure systems to a changing climate is critical for national security. Climate induced recurrent flooding is already impacting over 20 U.S. Air Force installations, underscoring the urgency of revisiting precipitation standards and stormwater infrastructure design. Despite growing scientific knowledge and an expanding set of tools for updating outdated precipitation standards based on the assumption of climate stationarity, the adoption of climate informed analyses remain limited in practice. This study utilizes an existing framework to update Intensity (or Depth)-Duration-Frequency (DDF) curves using an ensemble of future climate projections. Change factors in precipitation estimates are derived and applied to six USAF installations across the U.S. The analysis is further extended to evaluate the implications of climate-informed DDFs on stormwater infrastructure performance and flood analysis at Tyndall AFB. Results indicate that the current design precipitation estimates are likely to become obsolete in all six USAF bases by the end of the century. The wide range of change factors across 32 GCM ensembles highlights the need to integrate uncertainty and evolving scientific data into infrastructure planning. The study also finds that the impacts of a changing climate vary spatially and temporally, emphasizing the value of localized analysis for infrastructure decision-making. The work advances ongoing DoD and societal efforts to implement adaptation strategies aimed at enhancing infrastructure resilience.

Intensity-duration-frequency curves↗

Reduced‐Order Probabilistic Emulation of Physics‐Based Ring Current Models: Application to RAM‐SCB Particle Flux

Abstract In this work, we address the computational challenge of large‐scale physics‐based simulation models for the ring current. Reduced computational cost allows for significantly faster than real‐time forecasting, enhancing our ability to predict and respond to dynamic changes in the ring current, valuable for space weather monitoring and mitigation efforts. Additionally, it can also be used for a comprehensive investigation of the system. Thus, we aim to create an emulator for the Ring current‐Atmosphere interactions Model with Self‐Consistent magnetic field (RAM‐SCB) particle flux that not only improves efficiency but also facilitates forecasting with reliable estimates of prediction uncertainties. The probabilistic emulator is built upon the methodology developed by Licata and Mehta (2023), https://doi.org/10.1029/2022sw003345 . A novel discrete sampling is used to identify 30 simulation periods over 20 years of solar and geomagnetic activity. Focusing on a subset of particle flux, we use Principal Component Analysis for dimensionality reduction and Long Short‐Term Memory (LSTM) neural networks to perform dynamic modeling. Hyperparameter space was explored extensively resulting in about 5% median symmetric accuracy across all data sets for one‐step dynamic prediction. Using a hierarchical ensemble of LSTMs, we have developed a reduced‐order probabilistic emulator (ROPE) tailored for time‐series forecasting of particle flux in the ring current. This ROPE offers accurate predictions of omnidirectional flux at a single energy with no pitch angle information, providing robust predictions on the test set with an error score below 11% and calibration scores under 8% with bias under 2% providing a significant speed up as compared to the full RAM‐SCB run.

79 ASTRONOMY AND ASTROPHYSICS↗

Revealing the evolution of order in materials microstructures using multi-modal computer vision

The development of high-performance materials for microelectronics, energy storage, and extreme environments depends on our ability to describe and direct property-defining microstructural order. Our present understanding is typically derived from laborious manual analysis of imaging and spectroscopy data, which is difficult to scale, challenging to reproduce, and lacks the ability to reveal latent associations needed for mechanistic models. Here, we demonstrate a multi-modal machine learning (ML) approach to describe order from electron microscopy analysis of the complex oxide La 1−x Sr x FeO 3 . We construct a hybrid pipeline based on fully and semi-supervised classification, allowing us to evaluate both the characteristics of each data modality and the value each modality adds to the ensemble. We observe distinct differences in the performance of uni- and multi-modal models, from which we draw general lessons in describing crystal order using computer vision.

36 MATERIALS SCIENCE↗

Toward machine-learning-assisted PW-class high-repetition-rate experiments with solid targets

We present progress in utilizing a machine learning (ML) assisted optimization framework to study the trends in a parameter space defined by spectrally shaped, high-intensity, petawatt-class (8 J, 45 fs) laser pulses interacting with solid targets and give the first simulation-based overview of predicted trends. A neural network (NN) incorporating uncertainty quantification is trained to predict the number of hot electrons generated by the laser–target interaction as a function of pulse shaping parameters. The predictions of this NN serve as the basis function for a Bayesian optimization framework to navigate this space. For post-experimental evaluation, we compare two separate neural network (NN) models. One is based solely on data from experiments, and the other is trained only on ensemble particle-in-cell simulations. Reviewing the predicted and observed trends across the experiment-capable laser parameter search space, we find that both ML models predict a maximal increase in hot electron generation at a level of approximately 12%–18%; however, no statistically significant enhancement was observed in experiments. On direct comparison of the NN models, the average discrepancy is 8.5%, with a maximum of 30%. Since shot-to-shot fluctuations in experiments affect the observations, we evaluate the behavior of our optimization framework by performing virtual experiments that vary the number of repeated observations and the noise levels. Here, we discuss the implications of such a framework for future autonomous exploration platforms in high-repetition-rate experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

PV Generation and Load Forecasting for Adjuntas PR Community Microgrids

Existing frameworks to forecast time-series photovoltaic (PV) output power and consumer load for microgrid operations and controls assume a near-continuous availability of real-time input features from the field assets such as PV inverters, energy meters, and weather station. These incoming data points are used to periodically retrain models and update forecast snapshots over a moving horizon window, be it one hour-ahead, one-day ahead, or one-week ahead. However, such frameworks are not resilient to disruptions in data availability caused by losses in communications between the field sensors and data loggers. Hence, there is a need for programs that assume no availability of real-time microgrid asset data and still make reliable forecasts that can be used for decision-making. Such programs would be apt to function in extreme weather events such as hurricanes and would use lightweight recursive time-series models to independently forecast solar irradiance and ambient temperature, then compute PV power from those forecasts, as well as independently forecast consumer load. The codebase performs forecasting for the scenario of when the microgrid does not have a reliable access to forecasts or real-time observations of solar irradiance (I) and ambient temperature (AT) and load (Load) to be able to adequately forecast, in real-time, the PV power production or a business' load. In this case, using historical values of PV power and load, a univariate forecasting of generation and consumption are respectively made. The use-case in particular has two sub-scenarios: one, a normal 7-day ahead forecast where the unavailability of real-time data is assumed due to infrastructure issues such as loss of communication or sensor maintenance or service downtimes. Whereas a hurricane-caused unavailability of real-time data requires a second model trained specifically on historical hurricane days to be able to capture the extreme day behavior of generation in particular, and load if applicable. A gradient boosted regression tree comprises an ensemble of additive models that map between the input of historical values (be it irradiance, temperature, or load) and their corresponding output forecasts of a given horizon such that the individual learner predictions are summed up over the total number of such learners in the ensemble to produce an aggregate forecast. A weighting mechanism is applied to the training data in each iteration, where actual and forecast values are compared to penalize incorrect forecasts by increasing the weight and reducing it to reward correct forecasts. The code's benefits are that it: (a) accounts for a contingency where communication loss renders newly measured real-time data unavailable for model tuning and snapshot updates; (b) presents blind forecasting that recursively determines the next time-step value in a horizon using the forecast of the same attribute from a prior step; and (c) employs lightweight models that, once trained, can reliably generalize for different horizons, which make them suitable for enhancing the resilience of field microgrids prone to extreme events that encounter disruptions to data availability.

Sundararajan, Aditya [Oak Ridge National Laborator↗

Theoretical Modeling for Numerical Weather Prediction

The goal is to utilize predictability theory and numerical experimentation to identify and understand some of the dynamical processes which must be modeled more realistically if large-scale numerical weather predictions are to be improved. The emphasis is on the use of relatively simple models to exlore the properties of physically comprehensive general circulation models (GCM's). A global linear quasi-geostrophic model and the Goddard Laboratory for Atmospheric Sciences (GLAS) GCM were used to investigate several mechanisms which are responsible for the decay of large-scale forecast skill in mid-latitude numerical weather predictions. Five-day forecasts for an ensemble of cases were made using First GARP Global Experiment data. It was found that forecast skill depends crucially on the specification of the stationary forcing. A lack of stationary forcing leads to spurious westwad propagation of the ultralong waves. Forecasts made with stationary forcings derived from climatological data are superior to those using forcings inferred from observations immediately preceding the forecast period. Interhemispheric forecast differences were analyzed, and the model errors were compared to errors of a simple persistence-damped-to-climatology scheme and to errors of the GLAS GCM.

Somerville, R. C. J.↗

Final shuttle-derived atmospheric database: Development and results from thirty-two flights

The final Shuttle-derived atmospheric data base is presented. The relational data base is comprised of data from 32 Space Transportation System (STS) descent flights, to include available meteorology data taken in support of each flight. For the most part, the available data are restricted to the middle atmosphere. In situ accelerations, sensed by the tri-redundant Inertial Measurement Unit (IMU) to an accuracy better than 1 mg, are combined with post-flight Best Estimate Trajectory (BET) information and predicted, flight-substantiated Orbiter aerodynamics to provide determinations up to altitudes of 95 km. In some instances, alternate accelerometry data with micro-g resolution were utilized to extend the data base well into the thermosphere. Though somewhat limited, the ensemble of flights permit a reasonable sampling of monthly, seasonal, and latitudinal variations which can be utilized for atmospheric science investigations and model evaluations and upgrades as appropriate. More significantly, the unparallel vertical resolution in the Shuttle-derived results indicate density shears normally associated with internal gravity waves or local atmospheric instabilities. Consequently, these atmospheres can also be used as stress-atmospheres for Guidance, Navigation and Control (GN and C) system development and analysis as part of any advanced space vehicle design activities.

Findlay, John T.↗

PIV Data Validation Software Package

A PIV data validation and post-processing software package was developed to provide semi-automated data validation and data reduction capabilities for Particle Image Velocimetry data sets. The software provides three primary capabilities including (1) removal of spurious vector data, (2) filtering, smoothing, and interpolating of PIV data, and (3) calculations of out-of-plane vorticity, ensemble statistics, and turbulence statistics information. The software runs on an IBM PC/AT host computer working either under Microsoft Windows 3.1 or Windows 95 operating systems.

Blackshire, James L.↗

Data-driven prediction of scaling and ignition of inertial confinement fusion experiments

Recent advances in inertial confinement fusion (ICF) at the National Ignition Facility (NIF), including ignition and energy gain, are enabled by a close coupling between experiments and high-fidelity simulations. Neither simulations nor experiments can fully constrain the behavior of ICF implosions on their own, meaning pre- and postshot simulation studies must incorporate experimental data to be reliable. Linking past data with simulations to make predictions for upcoming designs and quantifying the uncertainty in those predictions has been an ongoing challenge in ICF research. We have developed a data-driven approach to prediction and uncertainty quantification that combines large ensembles of simulations with Bayesian inference and deep learning. The approach builds a predictive model for the statistical distribution of key performance parameters, which is jointly informed by past experiments and physics simulations. The prediction distribution captures the impact of experimental uncertainty, expert priors, design changes, and shot-to-shot variations. We have used this new capability to predict a 10× increase in ignition probability between Hybrid-E shots driven with 2.05 MJ compared to 1.9 MJ, and validated our predictions against subsequent experiments. We describe our new Bayesian postshot and prediction capabilities, discuss their application to NIF ignition and validate the results, and finally investigate the impact of data sparsity on our prediction results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗