Search NASA⌕ Search

SEARCH · Search NASA

Results for “data model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

From Machine Learning to Machine Reasoning: A Model-based Approach to Analyze Equipment Reliability Data

In current nuclear power plants (NPPs) a large amount of condition-based data which can be used to assess and monitor component health and performance. Assessing component health from such data can be performed with a large variety of methods. While the analysis of numeric data can be performed with several methods, the extraction of information from textual data remains a challenge. Currently employed natural language processing (NLP) methods do not really provide quantitative information that might be contained in IRs. In addition, the integration of numeric and textual data to identify possible causal relationships between data elements is still an unresolved challenge. This paper presents an approach to extract information from textual (e.g., incident or maintenance reports) and numeric data that relies on model based system engineer (MBSE) models. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence while semantic analysis is designed to analyze the logic structure of a sentence. An innovative element of our approach is that semantic analysis uses MBSE models to identify links between textual elements. Similarly, numeric data is directly linked to elements of the MBSE models in order to map which functions are being monitored.

97 - MATHEMATICS AND COMPUTING↗

Neural Posterior Estimation for Scalable and Accurate Inverse Parameter Inference in Li-Ion Batteries

Diagnosing the internal state of Li-ion batteries is critical for battery research, operation of real-world systems, and prognostic evaluation of remaining lifetime. By using physics-based models to perform probabilistic parameter estimation via Bayesian calibration, diagnostics can account for the uncertainty due to model fitness, data noise, and the observability of any given parameter. However, Bayesian calibration in Li-ion batteries using electrochemical data is computationally intensive even when using a fast surrogate in place of physics-based models, requiring many thousands of model evaluations. A fully amortized alternative is neural posterior estimation (NPE). NPE shifts the computational burden from the parameter estimation step to data generation and model training, reducing the parameter estimation time from minutes to milliseconds, enabling real-time applications. The present work shows that NPE can infer parameters equally or more accurately than Bayesian calibration, even if it leads to higher voltage reconstruction errors. We also demonstrate that the higher computational costs for data generation are tractable even in high-dimensional cases (ranging from 6 to 27 estimated parameters). The NPE method also offers several interpretability advantages over Bayesian calibration, such as local parameter sensitivity to specific regions of the voltage curve. The NPE method is demonstrated using an experimental fast charge dataset, with parameter estimates validated against measurements of loss of lithium inventory and loss of active material. The implementation is made available in a companion repository (https://github.com/NatLabRockies/BatFIT).

25 ENERGY STORAGE↗

Data for KETCHUP: Parameterizing of Large-Scale Kinetic Models Using Multiple Datasets with Different Reference States

Repository for Kinetic Estimation Tool Capturing Heterogeneous Datasets Using Pyomo (KETCHUP), a flexible parameter estimation tool that leverages a primal-dual interior-point algorithm to solve a nonlinear programming (NLP) problem that identifies a set of parameters capable of recapitulating the steady-state fluxes and concentrations in wild-type and perturbed metabolic networks. KETCHUP can use K-FIT [2] input files. Example K-FIT input files are located in the K-FIT repository at https://github.com/maranasgroup/K-FIT.

Metabolomics↗

Anomaly Detection In DUNE Using AI/ML

We designed and built an AI/ML model to detect anomalies in the DUNE far detector data. The model has been trained on simulated radiological background (rbkg) data, which is the major background for supernova burst neutrinos. The trained model was evaluated on both new samples of radiological backgrounds and supernova burst neutrino events in the elastic scattering and charged current interaction channels. We found that the trained model can successfully identify supernova burst neutrino events as anomalies while identifying radiological backgrounds as nominal events.

Novello, Eric [Unlisted, US]↗

PIPES (Pipeline for Integrated Projects in Energy Systems) [SWR-24-89]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. https://github.com/nrel-pipes/pipes-api https://github.com/nrel-pipes/pipes-web https://github.com/nrel-pipes/nrel-pipes

Gu, Jianli↗

Optimal Control of an Oscillating Surge Wave Energy Converter

During this project, we experimentally investigated the hydrodynamics and performance of a laboratory-scale oscillating surge wave energy converter (OSWEC).We looked at how flap buoyancy and driveline losses (primarily in the form of stiction) affected the dynamics and performance of the device. In addition, we assessed the influence of flap profile (rounded vs. square edges) on OSWEC hydrodynamics. Through this, we were able to develop a deeper understanding of OSWEC performance and provide guidance on strategies to counteract artifacts that may be present in laboratory models, but are absent in field-scale devices. To do this, we tested a laboratory-scale OSWEC in the Sea Wave Environmental Lab (SWEL) wave tank at the National Renewable Energy Laboratory (NREL). We ran several types of experiments to investigate the hydrodynamics and performance of the device. Overall, we achieved the overall goal of experimentally investigating the hydrodynamics and performance of this device. We discovered important and unexpected trends in performance, and collected time-resolved data to help us further investigate the underlying hydrodynamics responsible for these trends. In addition, we are currently using the time-resolved data from these experiments to build data-driven models of the dynamics, which can in turn be used to inform data-driven model predictive control of this device and address this objective in the future.

16 TIDAL AND WAVE POWER↗

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING↗

Data and scripts associated with a manuscript modeling microbial regulation of priming effects

This data package is associated with the publication “Modeling Microbial Regulatory Feedback in Organic Matter Decomposition Identifies Copiotrophic Traits as Key Drivers of Positive Priming” published as a preprint on BioRXiv by Ahamed et al. (2026); https://doi.org/10.1101/2024.08.11.607483. The package contains MATLAB scripts and saved simulation outputs used to implement a cybernetic model of microbial regulation during complex organic matter (OM) decomposition governing priming effects. It includes models of (i) single microbial functional groups (copiotrophic or oligotrophic degraders) and (ii) binary consortia composed of degraders and non-degraders with contrasting or common growth traits. Simulation results were generated using Monte Carlo analyses, with randomized key model parameters across a range of environmental mixing fractions of complex and labile OM. The dataset was created to provide a transparent and reusable computational framework for systematically exploring how microbial growth traits, metabolic regulation, and community composition influence OM decomposition dynamics and priming effects. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes the variable definitions. This package includes: (1) annotated MATLAB code implementing the system of ordinary differential equations and cybernetic control laws; (2) saved output files containing data (e.g., biomass, substrates, enzyme levels, priming metrics); and (3) scripts for processing saved outputs and regenerating figures. Specifically, the data package contains three main MATLAB scripts: runPrimingModel.m, runPlotData.m, and runPlotSuppFigS1.m, along with this readme and supporting documentation. Users should begin with runPrimingModel.m, which contains the annotated code implementing the system of ordinary differential equations and cybernetic control laws. This script runs the Monte Carlo simulations of microbial OM decomposition and allows users to modify microbial trait definitions, adjust parameter distributions, or define new community configurations. Simulation outputs are automatically saved as .mat files in the folder named SavedData, which stores all pre-generated results included in this package. The second script, runPlotData.m, reads files from the SavedData folder and processes them to regenerate the figures presented in the manuscript. The third script, runPlotSuppFigS1.m, specifically generates Figure S1 in the Supplementary Material of the manuscript. The package also includes the aforementioned files in non-proprietary .txt format. If users intend to use them, they should first save the files in their respective .m or .mat formats prior to execution in MATLAB.

Biomass concentration↗

Functional-type modeling approach and data-driven parameterization of methane emissions in wetlands (Final Technical Science Report)

Our goals are to improve understanding and quantitative representation of the multiple processes that affect methane emissions at a high (patch level, vertically detailed) spatial resolution, and translate this understanding to improved modeling capability of coastal wetland fluxes using the E3SM Land Model (ELM v1) wetland CH4 biogeochemistry module. We propose an experimental approach to identify and parameterize uncertainties in ELM. Understanding of methane emissions can be improved along three conceptual axes: (i) horizontal (ecohydrological patch resolution), (ii) vertical (through the depth of the soil column), and (iii) process level (e.g., resolving microbial pathways, vegetation specific transport pathways). Along each of the three axes, we will characterize, quantify, and model, the key ecological, hydrological, and meteorological controls of methane (CH4) flux heterogeneity in four model coastal wetlands.

54 ENVIRONMENTAL SCIENCES↗

Modeling light signals using data from the first pulsed neutron source program at the DUNE vertical drift ColdBox test facility at the CERN Neutrino Platform

In this paper, we present a first quantitative test of detected light signals produced in a pulsed neutron source run in a small vertical drift LArTPC at the CERN Neutrino Platform ColdBox test facility. The ColdBox cryostat, detectors, neutron sources, and particle interactions are modeled and simulated using Fluka. We demonstrate the ability to identify the contribution from neutron interactions using X-ARAPUCA photodetectors, and show first comparisons of data to simulation, which indicate reasonable agreement. A time constant is also fitted from the neutron-beam-off light signal spectrum and found consistent between data and simulation. Several important systematic effects are discussed and serve as guides for future runs at larger LArTPCs.

Detector modelling and simulations I (interaction ↗

Bias Correction and Statistical Downscaling of Solar Radiation Using NA-CORDEX and the NSRDB

The current state-of-art for estimating long-term PV production uses long-term estimates of solar radiation variables, such as global horizontal irradiance (GHI), from previous years. This data is used in models such as the System Advisor Model (SAM) or PYSyst to predict annual production for a PV plant. This information is then used to estimate the production over the next 20 years (a typical plant lifetime) under the assumption that the variability over the current period is representative of the future. As the PV industry moves to extend plant lifetimes to 50 years the current assumptions of representativeness of weather may not be appropriate. This is especially true as our climate changes rapidly. To assess long-term PV production, future projections for solar radiation based on projected carbon emissions are readily available in regional and global climate models. However, climate model projections contain inherent biases that may need to be corrected for accurate analysis of future projections of climate variables. Several studies have analyzed projections of solar radiation for future years, however the accuracy of the model output compared to current and historic data has not been widely studied. Chen (2021) showed that available climate models do not accurately represent solar radiation in some cases, over-projecting GHI at the surface while under-projecting its obstructions, such as clouds and aerosols. This works aims to (1) increase understanding of the accuracy of solar radiation currently available in global and regional climate models and (2) implement bias correction through linear models based on reanalysis data compared to observed solar radiation. The latter aim will be conducted using available observed solar radiation data and modeled data from several regional climate models (RCMs). The bias correction method will be applied to projections of solar radiation resulting in a more accurate representation of the future of solar production.

climate data↗

Integrating Characteristic Arctic Vegetation in a Land Surface Model Improves Representation of Carbon Dynamics Across a Tundra Landscape: Modeling Archive

This modeling archive is in support of the Next-Generation Ecosystem Experiments in the Arctic (NGEE Arctic) publication "Integrating Characteristic Arctic Vegetation in a Land Surface Model Improves Representation of Carbon Dynamics Across a Tundra Landscape", by Murphy et al. (2025). This archive contains model input files and outputs from landscape-scale simulations conducted using ELM, the land model component of the Department of Energy’s Energy Exascale Earth System Model (E3SM), at the Council NGEE Arctic field site (Council Road mile marker 71) on Alaska’s Seward Peninsula. Input data and model output from two sets of ELM simulations are provided. The first set of simulations were conducted with the two default ELM Arctic plant functional types (PFTs; broadleaf deciduous boreal shrub and a C3 grass) and the second set of simulations were conducted with a set of nine Arctic-specific PFTs including nonvascular mosses and lichens, graminoids, forbs, evergreen dwarf shrubs, three height classes of deciduous shrubs (dwarf, low, and low to tall), and deciduous alder shrubs (Sulman et al., 2021). Parameter names and major parameter changes in the Arctic-specific PFT configuration are described in Sulman et al. (2021) and archived in the Sulman et al. (2021) dataset (see below). Simulations were spatially explicit, covering an approximately 6.4X3.3 km domain at the Council site with a spatial resolution of 100 m for a total of 2,112 simulated grid cells under each ELM PFT configuration. The modeling archive contains meteorological forcing (seven *.nc files and one *.txt file), a domain definition file (one *.nc files), land surface configuration files (two *.nc files), parameter files (two *.nc files), annual ELM output files spanning 1980-2014 (68 *.nc files), and a User’s Guide (*pdf file). Additional information on the provided files is in the “Modeling Archive Contents” section of the User’s Guide. Model outputs are aggregated to the column scale (i.e. PFT-specific outputs are not provided here).

Murphy, Bailey [ORNL] (ORCID:0000000203995221)↗

A Survey of Open-Source Tools for Transmission and Distribution Systems Research

This work presents a review of open-source electric power transmission and distribution systems analysis tools suitable for use by industry professionals and academic researchers. Due to the high complexity of the electric grid, there exist numerous tools and extensive research pertaining to nearly every aspect of the design, operation, and control of transmission and distribution networks. In addition to the commercial tools, a wide range of free, open-source tools, models, and data usable by the scientific community for related research have been developed by different organizations, including both international and US universities and national laboratories. However, due to the absence of a catalog of available tools, models and data, researchers often lack a knowledge of existing capabilities and may develop duplicative software and tools. Increasing awareness of these available resources seeks to accelerate their broader use, leading to more efficient and standardized grid analysis. This review paper (which is part of a larger survey effort that studied over 400 tools in the transmission, distribution, buildings, and electric vehicles space) outlines selected open-source resources that have been developed in power transmission and distribution systems research. It is anticipated that this work can serve as a guide for industry and academic researchers alike, ensuring that research efforts are well-channeled.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Anomaly Detection in DUNE FD LArTPC Readouts

We designed and built an AI/ML model to detect anomalies in the DUNE far detector data. The model has been trained on simulated radiological background (rbkg) data, which is the major background for supernova burst neutrinos that we want to detect as anomaly in thios work. The trained model was evaluated on both new samples of radiological backgrounds and supernova burst neutrino events in the elastic scattering and charged current interaction channels. We found that the trained model can successfully identify supernova burst neutrino events as anomalies while identifying radiological backgrounds as nominal events.

Novello, Eric [Fermilab]↗

Guiding Principles for Geochemical/Thermodynamic Model Development and Validation in Nuclear Waste Disposal: A Close Examination of Recent Thermodynamic Models for H + —Nd 3+ —NO 3 - (—Oxalate) Systems

Development of a defensible source-term model (STM), usually a thermodynamical model for radionuclide solubility calculations, is critical to a performance assessment (PA) of a geologic repository for nuclear waste disposal. Such a model is generally subjected to rigorous regulatory scrutiny. In this article, we highlight key guiding principles for STM model development and validation in nuclear waste management. We illustrate these principles by closely examining three recently developed thermodynamic models with the Pitzer formulism for aqueous H + —Nd 3+ —NO 3 - (—oxalate) systems in a reverse alphabetical order of the authors: the XW model developed by Xiong and Wang, the OWC model developed by Oakes et al., and the GLC model developed by Guignot et al., among which the XW model deals with trace activity coefficients for Nd(III), while the OWC and GLC models are for concentrated Nd(NO 3 ) 3 electrolyte solutions. The principles highlighted include the following: (1) Principle 1. Validation against independent experimental data: A model should be validated against experimental data or field observations that have not been used in the original model parameterization. We tested the XW model against multiple independent experimental data sets including electromotive force (EMF), solubility, water vapor, and water activity measurements. The results show that the XW model is accurate and valid for its intended use for predicting trace activity coefficients and therefore Nd solubility in repository environments. (2) Principle 2. Testing for relevant and sensitive variables: Solution pH is such a variable for an STM and easily acquirable. All three models are checked for their ability to predict pH conditions in Nd(NO 3 ) 3 electrolyte solutions. The OWC model fails to provide a reasonable estimate for solution pH conditions, thus casting serious doubt on its validity for a source-term calculation. In contrast, both the XW and GLC models predict close-to-neutral pH values, in agreement with experimental measurements. (3) Principle 3. Honoring physical constraints: Upon close examination, it is found that the Nd(III)-NO 3 association schema in the OWC model suffers from two shortcomings. Firstly, its second stepwise stability constant for Nd(NO 3 ) 2+ (log K 2 ) is much higher than the first stepwise stability constant for NdNO 3 2+ (log K 1 ), thus violating the general rule of (log K 2 –log K 1 ) < 0, or $\frac{K1}{K2}$>1. Secondly, the OWC model predicts abnormally high activity coefficients for Nd(NO 3 ) 2 + (up to ~900) as the concentration increases. (4) Principle 4. Minimizing degrees of freedom for model fitting: The OWC model with nine fitted parameters is compared with the GLC model with five fitted parameters, as both models apply to the concentrated region for Nd(NO 3 ) 3 electrolyte solutions. The latter appears superior to the former because the latter can fit osmotic coefficient data equally well with fewer model parameters. The work presented here thus illustrates the salient points of geochemical model development, selection, and validation in nuclear waste management.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Data‐driven variational method for discrepancy modeling: Dynamics with small‐strain nonlinear elasticity and viscoelasticity

Abstract The effective inclusion of a priori knowledge when embedding known data in physics‐based models of dynamical systems can ensure that the reconstructed model respects physical principles, while simultaneously improving the accuracy of the solution in the previously unseen regions of state space. This paper presents a physics‐constrained data‐driven discrepancy modeling method that variationally embeds known data in the modeling framework. The hierarchical structure of the method yields fine scale variational equations that facilitate the derivation of residuals which are comprised of the first‐principles theory and sensor‐based data from the dynamical system. The embedding of the sensor data via residual terms leads to discrepancy‐informed closure models that yield a method which is driven not only by boundary and initial conditions, but also by measurements that are taken at only a few observation points in the target system. Specifically, the data‐embedding term serves as residual‐based least‐squares loss function, thus retaining variational consistency. Another important relation arises from the interpretation of the stabilization tensor as a kernel function, thereby incorporating a priori knowledge of the problem and adding computational intelligence to the modeling framework. Numerical test cases show that when known data is taken into account, the data driven variational (DDV) method can correctly predict the system response in the presence of several types of discrepancies. Specifically, the damped solution and correct energy time histories are recovered by including known data in the undamped situation. Morlet wavelet analyses reveal that the surrogate problem with embedded data recovers the fundamental frequency band of the target system. The enhanced stability and accuracy of the DDV method is manifested via reconstructed displacement and velocity fields that yield time histories of strain and kinetic energies which match the target systems. The proposed DDV method also serves as a procedure for restoring eigenvalues and eigenvectors of a deficient dynamical system when known data is taken into account, as shown in the numerical test cases presented here.

Masud, Arif↗

TxDOT Road Elevation Model Dataset

This dataset provides three formats of Road Elevation Model (REM) data: 3D road line/polygon GeoPackage (GPKG), road lidar LAZ and COPC LAZ, and road digital surface model (DSM) GeoTIFF. Data are produced from the ~50TB TxGIO (formerly TNRIS) state lidar collections. This dataset is currently organized by maintenance section in each TxDOT district. Computation is done on GPU computing resources at Oak Ridge National Laboratory (ORNL), through a Strategic Partnership Project with UT Austin and an NSF ACCESS computing allocation award that enables fast massive data movement between TACC Corral and ORNL CADES/OLCF using Globus. In addition to this release from ORNL, a copy of this dataset can also be downloaded at https://web.corral.tacc.utexas.edu/nfiedata/road3d/.

13 HYDRO ENERGY↗

Quantifying mean, variability, and uncertainty in indoor radon exposure in Pennsylvania using random forest and quantile regression forest models

Radon is a naturally occurring radioactive gas that poses a serious health risk as the primary cause of lung cancer in non-smokers. Despite the well-known adverse association with health outcomes, current radon exposure assessments are limited to county-level or average-level estimates, which fail to capture regional variability. This study uses Machine Learning models, including Random Forest (RF) and Quantile Regression Forest (QRF), to estimate the indoor radon concentrations at the ZCTA (Zip code tabulation area)-level and characterize uncertainties in model estimates. Incorporating geological, meteorological, and building-specific data, the models aim to improve radon risk assessment by capturing mean exposure, variability, and extreme concentration levels. Processed radon test data (n = 718,111) were analyzed using average, variability, and quantile prediction methods. Models that estimate the average radon exposure at the ZCTA-level can yield promising model-fit results, but they do not capture the underlying variability of indoor radon exposure within a ZCTA. We utilize volatility analyses to identify characteristics indicative of high variability of indoor radon exposure. We also show that a QRF model can be used to estimate upper quantiles of residential radon exposure, thereby uncovering localized areas of elevated exposure that were not apparent in mean estimates. The results highlighted the need for a deep characterization of exposure risk and show that regions with moderate average exposure levels could still harbor extreme outliers with implications for evaluating health risks. Utilizing multiple radon exposure models allows for a deeper characterization of radon risk within a geographic area and can better identify high-risk areas. The results from this study provide a foundation for developing mitigation strategies and examining associations between radon exposure and health outcomes at fine scales. Future research should extend the geographic scope and incorporate additional environmental risk factors to establish a comprehensive framework for risk assessment.

Lee, Heechan [ORNL]↗