Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Governance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Addressing Investment Barriers by Improving Documentation of Sustainable Biomass Resources (Workshop Report)

On May 8, 2025, Oak Ridge National Laboratory (ORNL), in collaboration with IEA Bioenergy Task 43 and the Biofuture Platform, convened an international workshop in Vancouver, Canada to improve the Global Biomass Resource Assessment. This effort addresses investment barriers in the global bioeconomy by improving the transparency, consistency, and usability of biomass supply data. The workshop gathered 38 participants from 11 countries, representing government agencies, academia, and industry. Participants reviewed the status of the biomass dataset, tested the data-sharing platform, and provided direct input on priorities for improvement.

09 BIOMASS FUELS↗

A Perspective on Traditional and Data Driven Electrochemical Modeling and Analysis

To understand the behavior of electrochemical systems, we need to reduce the dimensionality of the measured current-voltage-time (I-V-t) data by fitting models, thus enabling us to analyze and compare the governing physics. Traditionally, the process for this is an 'expert first' approach: defining the model and its explicit assumptions based on inductive reasoning or empirical observation, fitting small portions of the I-V-t data where assumptions are most valid or carefully designing experiments to enforce key assumptions, and then interpreting the model parameters. However, modern data-driven methods enable a new paradigm: a 'data first' approach, where the latent behaviors governing the system's measured response are identified directly using machine-learning models that optimize both model structure and parameters from the I-V-t data, guaranteeing that the learned model explains as much of the observed system response as possible. After model identification, the model can then be interrogated by an expert to connect observed behaviors with underlying physics. This talk will review several different types of electrochemical analysis (electrochemical impedance, differential voltage-capacity, electrochemical kinetics) and compare the traditional and data-driven methods for analyzing the data.

42 ENGINEERING↗

Data and scripts associated with the manuscript "Organic Molecules are Deterministically Assembled in River Sediments"

This data package is associated with the publication "Organic Molecules are Deterministically Assembled in River Sediments" submitted to Scientific Reports (Stegen et al., 2024). The study applies community ecology methods to dissolved organic matter (DOM) chemistry from variably inundated riverbed sediments to uncover principles governing DOM composition at a reach-scale. This data package documents the workflow used to process and generate the main findings in the manuscript. The R scripts reference the raw, unprocessed Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data from another data package, available on ESS-DIVE at https://data.ess-dive.lbl.gov/view/doi:10.15485/1834208. The scripts then process the raw FTICR-MS data and generate the findings and figures presented in the associated manuscript. In brief, this study demonstrates that DOM assemblages in variably inundated sediments are primarily governed by deterministic variable selection, including sediment moisture effecting the degree of deterministic assembly. See the manuscript for more details pertaining to interpretation and implications of the findings. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/ECA_2020_Sed.This data package is comprised of 6 scripts and 7 folders. The file-level metadata file (file ending in "flmd.csv") lists all files contained in this data package and descriptions for each. The data dictionary (file ending in "dd.csv) describes all tabular data columns and their respective definitions and units. The FTICR_Processing_Scripts produce the outputs found in the "Processed_Data" folder. The remaining scripts (located in the parent directory) produce the outputs found in the following four folders: (1) "MCD_Dendrograms", "MCD_Randomizations", "MCD_bNTI_Outcomes", and "OM_Null_Modeling". The fifth script additionally takes the three comma-separated values (CSV) files found in the parent directory as input ("VGC_texture.csv", "merged_weights.csv", and "ECA2_FTICR_BetaDisp.csv"). The outputs of each of the five scripts serve as the input to the following script, with the final outputs stored in the folder "OM_Null_Modeling".

54 ENVIRONMENTAL SCIENCES↗

Harnessing Satellite Data Alone for Mapping Global Thermal Anisotropy

Mapping thermal anisotropy across global lands is critical for advancing a wide range of Earth science studies. However, a comprehensive understanding of global thermal anisotropy intensity (TAI) and its governing factors remains missing. We introduce a novel data-driven methodology to quantify global TAI exclusively using multi-angle MODIS land surface temperature time series observations. Our analysis reveals distinct seasonal and diurnal TAI patterns, with global mean summertime TAI exceeding 2.9°C. Furthermore, we identify strong associations between TAI and key surface and atmospheric parameters, such as leaf area index and downward shortwave radiation. Our findings advocate for a paradigm shift from model-based to data-driven approaches in correcting thermal anisotropy, thereby addressing a critical bottleneck in Earth observation.

54 ENVIRONMENTAL SCIENCES↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

FY25 Report on Water NSTF Testing: Parametric and Accident Testing with Lower Tank Inlet

The Natural Convection Shutdown Heat Removal Test Facility (NSTF) at Argonne National Laboratory has continued to generate empirical validation data on the performance of water-based reactor cavity cooling system (RCCS) for seven years. This data is actively being used to support the development of passive decay heat removal systems for advanced reactors. Distinguishing this facility are 1) the large, ½ scale of the facility and 2) its governance under an NQA-1 qualified program for producing data of the highest pedigree for advanced reactor designers and regulators. In addition to the experimental activities discussed in this report, a computational modeling program continues to support the experimental program and further accuracy and understanding of the computational models. Together, the experimental and computational work create a mutually beneficial relationship integral to the overall program objective of advancing the understanding of the RCCS technology.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Filling the Gaps: A Bayesian Mixture Model for Imputing Missing Soil Water Content Data

ABSTRACT Soil water content (SWC) data are central to evaluating how soil moisture varies over time and space and influences critical plant and ecosystem functions, especially in water‐limited drylands. However, sensors that record SWC at high frequencies often malfunction, leading to incomplete timeseries and limiting our understanding of dryland ecosystem dynamics. We developed an analytical approach to impute missing SWC data, which we tested at six eddy flux tower sites along an elevation gradient in the southwestern United States. We impute missing data as a mixture of linearly interpolated SWC between the observed endpoints of a missing data gap and SWC simulated by an ecosystem water balance model (SOILWAT2). Within a Bayesian framework, we allowed the relative utility (mixture weight) of each component (linearly interpolated vs. SOILWAT2) to vary by depth, site and gap characteristics. We explored “fixed” weights versus “dynamic” weights that vary as a function of cumulative precipitation, average temperature, and time since the start of the gap. Both models estimated missing SWC data well ( R 2 = 0.70–0.88 vs. 0.75–0.91 for fixed vs. dynamic weights, respectively), but the utility of linearly interpolated versus SOILWAT2 values depended on site and depth. SOILWAT2 was more useful for more arid sites, shallower depths, longer and warmer gaps and gaps that received greater precipitation. Overall, the mixture model reliably gap‐fills SWC, while lending insight into processes governing SWC dynamics. This approach to impute missing data could be adapted to accommodate more than two mixture components and other types of environmental timeseries.

Ogle, Kiona [School of Informatics, Computing, and↗

Spatial analysis of cell patterning to aid genetic and phenotypic understanding of grass stomatal density: A case study in maize

Biological processes involve complex hierarchies where composite traits result from multiple component traits. However, holistically understanding of how sets of component traits interact to underpin genotype-to-phenotype relationships is generally lacking. Stomatal density (SD) is a tractable model system for exploring how high-throughput phenotyping (HTP) data could be exploited by a new spatial analysis approach to better understand a developmentally and functionally important trait. SD is a composite trait, resulting from various components related to cell identity and size, which are themselves governed by a series of spatio-developmental processes. Data from 192 recombinant inbred lines of maize [Zea mays (L.)] were analyzed by a new stomatal patterning phenotype (SPP) to (1) describe the average spatial probability distribution of the nearest neighboring stomata; (2) derive a core set of component traits related to cell size, cell packing, and positional probabilities; (3) build a structural equation model of component traits underlying SD; and (4) identify stomatal patterning quantitative trait loci (QTL). The core set of SPP-derived traits explained 74% of the variation in SD. Analyzing SPP component traits allowed some loci previously identified as generic SD QTL to be recognized as specific to lateral versus longitudinal elements of stomatal patterning. Therefore, this study highlights how novel insights can be gained by decomposing a composite trait (e.g., SD) into a set of component traits that were present in HTP data but not previously exploited.

59 BASIC BIOLOGICAL SCIENCES↗

Utah FORGE 5-2615: Laboratory Data for Insights on Hydraulic Fracture Closure and Stress Measurement

This dataset includes data from injection/fall-off experiments conducted in controlled laboratory settings. The aim is to investigate the physics governing fracture closure and the associated stress measurements during hydraulic fracturing. These time series data include flow rate, pressure, and volume measurements. These experiments were conducted as part of Utah FORGE Project 5-2615: Thermo-poromechanical Response of Fractured Rock.

15 GEOTHERMAL ENERGY↗

Rapid Inverse Parameter Inference Using Physics-Informed Neural Network

As Li-ion batteries become more essential in today's economy, tools need to be developed to accurately and rapidly diagnose a battery's internal state-of-health. Using a Li-ion battery's (high-rate) voltage response, it is proposed to determine a battery's internal state through Bayesian calibration. However, Bayesian calibration is notoriously slow and requires thousands of model runs. To accelerate parameter inference using Bayesian calibration, a surrogate model is developed to replace the underlying physics-based Li-ion model. Developing a surrogate model for rapid Bayesian calibration analysis is discussed for both the single particle model (SPM) and the pseudo two-dimensional (P2D) model. Surrogate models are constructed using physics-informed neural networks (PINNs) that encode the influence of internal properties on observed voltage responses. In practice, a neural network can be trained by: 1) using simulation results of the physics-based model (i.e., a data-loss approach); 2) using the residuals of the governing equations themselves (i.e., a physics-loss approach); or 3) using a combination of simulation results and governing equation residuals. In the present work, PINNs are developed using a variety of training losses and neural network architectures. In this analysis, it is shown that a PINN surrogate model can be reliably trained with only physics-informed loss. However, using a coupled data-informed and physics-loss approach produced the most accurate PINNs.

Bayesian calibration↗

Learning Nonlinear Reduced Models from Data with Operator Inference

This review discusses Operator Inference, a nonintrusive reduced modeling approach that incorporates physical governing equations by defining a structured polynomial form for the reduced model, and then learns the corresponding reduced operators from simulated training data. The polynomial model form of Operator Inference is sufficiently expressive to cover a wide range of nonlinear dynamics found in fluid mechanics and other fields of science and engineering, while still providing efficient reduced model computations. The learning steps of Operator Inference are rooted in classical projection-based model reduction; thus, some of the rich theory of model reduction can be applied to models learned with Operator Inference. This connection to projection-based model reduction theory offers a pathway toward deriving error estimates and gaining insights to improve predictions. Furthermore, through formulations of Operator Inference that preserve Hamiltonian and other structures, important physical properties such as energy conservation can be guaranteed in the predictions of the reduced model beyond the training horizon. This review illustrates key computational steps of Operator Inference through a large-scale combustion example.

Mechanics↗

PINN surrogate of Li-ion battery models for parameter inference, Part II: Regularization and application of the pseudo-2D model

Bayesian parameter inference is useful to improve Li-ion battery diagnostics and can help formulate battery aging models. However, it is computationally intensive and cannot be easily repeated for multiple cycles, multiple operating conditions, or multiple replicate cells. To reduce the computational cost of Bayesian calibration, numerical solvers for physics-based models can be replaced with faster surrogates. A physics-informed neural network (PINN) is developed as a surrogate for the pseudo-2D (P2D) battery model calibration. For the P2D surrogate, additional training regularization was needed as compared to the PINN single-particle model (SPM) developed in Part I. Both the PINN SPM and P2D surrogate models are exercised for parameter inference and compared to data obtained from a direct numerical solution of the governing equations. A parameter inference study highlights the ability to use these PINNs to calibrate scaling parameters for the cathode Li diffusion and the anode exchange current density. By realizing computational speed-ups of ~2250x for the P2D model, as compared to using standard integrating methods, the PINN surrogates enable rapid state-of-health diagnostics. Finally, in the low-data availability scenario, the testing error was estimated to ~2 mV for the SPM surrogate and ~10 mV for the P2D surrogate which could be mitigated with additional data.

25 ENERGY STORAGE↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

The AEOLUS Center is dedicated to developing a unified optimization-under-uncertainty framework for (1) learning predictive models from data and (2) optimizing experiments, processes, and designs governed by these models, all driven by complex, uncertain energy systems. AEOLUS addressed the critical need for principled, rigorous, scalable, and structure-exploiting capabilities for exploring parameter and decision spaces of complex forward simulation models---the so-called outer loop. This report summarizes the work done under DE-SC0021077 on (1) nonlocal models for solidification problems, (2) a multifidelity method for a nonlocal diffusion model, and (3) multifidelity Monte Carlo methods.

97 MATHEMATICS AND COMPUTING↗

Illinois Storage Corridor CarbonSAFE Phase III: Stakeholder Engagement and Outreach Plan

The Stakeholder Engagement and Outreach Plan provides a comprehensive framework for engaging stakeholders of the Illinois Storage Corridor (ISC) project. The ISC project is a CarbonSAFE Phase III project designed to facilitate commercial deployment of carbon capture, utilization, and storage (CCUS) in Illinois. The project aims to establish a multi-industry carbon storage corridor through development of storage sites near the One Earth Energy (OEE) ethanol production facility in north-central Illinois and the Prairie State Generating Company (PSGC) coal-fired power plant in south-central Illinois, with combined annual CO 2 capture ultimately exceeding 8.6 million tons per year. Stakeholder engagement is recognized as a critical component for successful CCUS deployment, alongside technical and economic considerations. As an emerging technology, CCUS may not be well understood by the general population, and lack of public awareness can lead to opposition that poses significant barriers to project development. This plan addresses this challenge through systematic stakeholder identification, analysis, planning, and implementation of engagement actions. The plan is structured around four main sections: Communication, Stakeholder Analysis, Stakeholder Engagement, and Environmental Justice. Activities will be conducted under Tasks 1 and 4 of the project's Statement of Project Objectives, with two key subtasks: (1) developing a stakeholder analysis and engagement plan through face-to-face meetings, facilitated discussions, and surveys; and (2) implementing stakeholder engagement and public outreach activities including meetings, open houses, and permit hearings. The Illinois State Geological Survey (ISGS) will manage engagement activities following DOE-NETL best practices, focusing on providing objective, fact-based information about CCUS and the ISC project. A comprehensive Communication Plan establishes protocols for media contacts, site visits, and crisis communications. The stakeholder analysis follows a structured workflow process divided into Pre-feasibility and Feasibility phases, incorporating contextual understanding, assessment, data collection, and analysis. Key stakeholder groups include government bodies, educational organizations, conservation and environmental groups, agricultural communities, and religious organizations. The plan addresses common stakeholder questions regarding project risks, benefits, safety, property values, liability, and environmental impacts. Recommendations emphasize developing clear messaging, creating informational materials, and preparing to address both project-specific and broader environmental concerns to ensure transparent communication and build stakeholder support throughout project implementation.

25 ENERGY STORAGE↗

Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems (Final Report for AEOLUS)

The AEOLUS Center is dedicated to developing a unified optimization-under-uncertainty framework for: (1) learning predictive models from data; and (2) optimizing experiments, processes, and designs governed by these models, all driven by complex, uncertain energy systems. AEOLUS addresses the critical need for principled, rigorous, scalable, and structure-exploiting capabilities for exploring parameter and decision spaces of complex forward simulation models. This report summarizes the key highlights of our research during the period of performance.

97 MATHEMATICS AND COMPUTING↗

CMIP7 Data Request: atmosphere priorities and opportunities

This paper presents a comprehensive overview of the Coupled Model Intercomparison Project Phase 7 (CMIP7) request for data unlocking key research avenues in atmospheric science and provides justification for the resources needed to produce this data. Topics within the CMIP7 Atmosphere Theme centre around processes and feedbacks in atmospheric science such as clouds, aerosols and atmospheric chemistry, atmospheric circulation, temperature variability and extremes, radiative forcings, and Earth system model evaluation. These topics are summarised in this paper as scientific “opportunities” which will be realised through CMIP7 experiments and Earth system model outputs. These opportunities were submitted by a thematic group of atmospheric science community representatives combined with an extended consultation process. The production of these variables will close key gaps and uncertainties identified during previous rounds of CMIP, and will be broadly used by scientific, policy, governmental, industry, and other communities that rely on climate model projections for research and decision making, including supporting the 7th Intergovernmental Panel on Climate Change Assessment Report (AR7). As an author group, we also reflect on the process used to collate this data request and make recommendations to future CMIP governance on implementing a consultation on this scale in the future.

58 GEOSCIENCES↗

A Statistician’s Overview of Physics-Informed Neural Networks for Spatio-Temporal Data

The recent success of deep neural network models with physical constraints (so-called, Physics-Informed Neural Networks, PINNs) has led to renewed interest in the incorporation of mechanistic information in predictive models. Statisticians and others have long been interested in this problem, which has led to several practical and innovative solutions dating back decades. In this overview, we focus on the problem of data-driven prediction and inference of dynamic spatio-temporal processes that include mechanistic information, such as would be available from partial differential equations, with a strong focus on the quantification of uncertainty associated with data, process, and parameters. Here, we give a brief review of several paradigms and focus our attention on Bayesian implementations given they naturally accommodate uncertainty quantification. We then show that it is straight-forward to include the Bayesian PINN (B-PINN) within the Bayesian hierarchical model (BHM) framework that has long been considered for modeling dynamic spatio-temporal processes. Such a BHM-PINN is illustrated via a simulation study in which a latent nonlinear Burgers’ equation PDE governs the dynamics of Poisson distributed spatio-temporal data. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.

Bayesian↗