Search NASA⌕ Search

SEARCH · Search NASA

Results for “data model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Data Assimilation with Machine Learning Surrogate Models: A Case Study with FourCastNet

Modern data-driven surrogate models for weather forecasting provide accurate short-term predictions but inaccurate and nonphysical long-term forecasts. This paper investigates online weather prediction using machine learning surrogates supplemented with partial and noisy observations. We empirically demonstrate and theoretically justify that, despite the long-time instability of the surrogates and the sparsity of the observations, filtering estimates can remain accurate in the long-time horizon. As a case study, we integrate the Fourier Forecasting Neural Network (FourCastNet), a weather surrogate model, within a variational data assimilation framework using partial, noisy ERA5 global reanalysis data from the European Centre for Medium-Range Weather Forecasts (ECMWF). Here, our results show that filtering estimates remain accurate over a year-long assimilation window and provide effective initial conditions for forecasting tasks, including extreme event prediction.

Data assimilation↗

Data Assimilation with Machine Learning Surrogate Models: A Case Study with FourCastNet

Modern data-driven surrogate models for weather forecasting provide accurate short-term predictions but inaccurate and nonphysical long-term forecasts. This paper investigates online weather prediction using machine learning surrogates supplemented with partial and noisy observations. We empirically demonstrate and theoretically justify that, despite the long-time instability of the surrogates and the sparsity of the observations, filtering estimates can remain accurate in the long-time horizon. As a case study, we integrate FourCastNet, a weather surrogate model, within a variational data assimilation framework using partial, noisy ERA5 data. Our results show that filtering estimates remain accurate over a year-long assimilation window and provide effective initial conditions for forecasting tasks, including extreme event prediction.

Adrian, Melissa [Univ. of Chicago, IL (United Stat↗

An Initial Microstructurally Informed Model of High Burnup Structure Formation in UO 2 Fuel

The microstructure of a UO 2 fuel pellet changes as burnup increases, impacting fuel performance. Predicting and characterizing high burnup structure (HBS) and dark zone formation is a key part of supporting burnup limit extensions for light water reactors. This paper describes a model developed through fitting radially resolved pellet data obtained from recently published microstructural characterization data. The model predicts grain size and grain character, in addition to pore density and size, with fitting dependencies on power history variables. Separately fitting power history variables to microstructural parameters allows for insight into the underlying physical phenomena for future model development. Additionally, experimental data have been correlated to an HBS fraction to facilitate the development of a model capable of predicting a total fuel restructured fraction at the engineering scale. In conclusion, this two-step approach provides a coupling from reactor power history to microstructural data to fractional HBS and creates a basis to model HBS-dependent parameters in a fuel performance code.

High burnup structure↗

Assessment of Storm-Associated Precipitation and Its Extremes Using Observational Data Sets and Climate Model Short-Range Hindcasts

Heavy precipitation, often associated with weather phenomena such as tropical cyclones, extratropical cyclones (ETCs), atmospheric rivers (ARs), and mesoscale convective systems (MCSs), can cause significant socio-economic loss. Here, in this study, we apply atmospheric feature trackers to quantify the contributions of these storm types in observational data sets and climate model short-range hindcasts. We generate a global hourly storm data set at 0.25° spatial resolution covering 2006–2020, based on the tracking results from TempestExtremes and Python FLEXible object TRacKeR. Our analyses show that these four storm types account for 67% of global annual mean precipitation and 82% of top 1% precipitation extremes, with MCSs mainly over the tropics, and ARs and ETCs over the midlatitudes. The percentage of precipitation contributions from these storms also show strong seasonality over many geographical locations. We further apply the tracking results to the Energy Exascale Earth System Model (E3SM) short-range hindcasts and evaluate how well these storms are simulated. The evaluation show that E3SM, with ∼1° resolution, significantly underestimates storm-associated precipitation totals and extremes, especially for MCSs in the tropics. Our analysis also suggests that model fails to capture the correct mean diurnal phases and amplitude of MCS precipitation. This phenomenon-based approach provides a better understanding of precipitation characteristics and can lead to enhanced model evaluation by revealing underlying problems in model physics related to precipitation processes associated with the heavy-precipitating storms.

54 ENVIRONMENTAL SCIENCES↗

Deep Learning-enhanced Block-Diagram Modeling of Solar Power Systems

Data-driven models of power system inverter-based resources are desired to run simulations faster than with detailed electromagnetic transient models, to hide proprietary design details, to support control system design applications, and to aggregate the effects of distributed energy resources. This paper applies a customized Hammerstein Wiener framework to train block diagram models from thousands of electromagnetic transient simulations or experimental test records. The block diagram models integrate with larger grid simulations as voltagecontrolled current sources or current-controlled voltage sources for several simulators. Guidelines for block architecture and training are presented. Three-phase balanced, three-phase unbalanced, and single-phase examples all achieve an acceptable root mean square error of no more than 0.05 per-unit.

Mcdermott, Thomas E. [Private consulting company]↗

Self‐Potential Tomography Preconditioned by Particle Swarm Optimization—Application to Monitoring Hyporheic Exchange in a Bedrock River

Abstract A self‐potential (SP) data‐inversion algorithm was developed and tested on an analytical model of electrical‐potential profile data attributed to single and multiple polarized electrical sources. The developed algorithm was then validated by an application to SP‐monitoring field data measured on the floodplain of East Fork Poplar Creek, Oak Ridge, Tennessee, to image electrical sources in areas conducive to preferential flow into the flood plain from the bedrock‐lined riverbed. The algorithm combined stochastic source‐localization by particle‐swarm‐optimization (PSO) of electrical sources characterized by simplified geometries with source tomography by regularized weighted least‐squares minimization of a quadratic objective function. Prior information was incorporated by preconditioning the tomography algorithm by PSO results. Variable percentages of random noise were added to analytical‐model data to evaluate the algorithm performance. Results indicated that true parameters of single‐source models were inverted and approximated with small residual error, whereas inversion of analytical‐model data representing multiple electrical sources accurately approximated the locations of the sources but miscalculated some parameters because of the non‐uniqueness of the inverse‐model solution. Source tomography applied to analytical model data during testing produced a spatially continuous parameter field that identified the locations of point‐scale synthetic dipole sources of electrical current flow with varying degrees of accuracy depending on the prior information incorporated into the tomography. When applied to SP‐monitoring field data, the algorithm imaged electrical sources within a known fault that intersects the bedrock riverbed and flood plain of East Fork Poplar Creek and depicted dynamic electrical conditions attributed to hyporheic exchange.

54 ENVIRONMENTAL SCIENCES↗

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O↗

Learning of networked spreading models from noisy and incomplete data

Recent years have seen a lot of progress in algorithms for learning parameters of spreading dynamics from both full and partial data. Some of the remaining challenges include model selection under the scenarios of unknown network structure, noisy data, missing observations in time, as well as an efficient incorporation of prior information to minimize the number of samples required for an accurate learning. Here, in this work, we introduce a universal learning method based on a scalable dynamic message-passing technique that addresses these challenges often encountered in real data. The algorithm leverages available prior knowledge on the model and on the data, and reconstructs both network structure and parameters of a spreading model. We show that a linear computational complexity of the method with the key model parameters makes the algorithm scalable to large network instances.

97 MATHEMATICS AND COMPUTING↗

Challenges and Opportunities for Electric Utility Modeling and Asset Valuation Frameworks: Case Study on Valuing New Pumped Storage Hydropower

Asset valuation by electric utilities is becoming increasingly difficult in the rapidly changing electric sector. Rapid deployment of variable generation and inverter-based storage systems along with uncertain demand growth, climate, policies, and other factors create a challenging environment for understanding the value proposition of a new potential asset. This report describes an effort between the Tennessee Valley Authority (TVA) and three U.S. Department of Energy laboratories to perform a detailed review of utility modeling and analysis practices for asset valuation and identify challenges and opportunities for advancing its methods into the future. It focuses on a case study of new potential pumped storage hydropower (PSH) because of growing interest in new PSH capacity to provide energy balancing, firm capacity, and a range of ancillary services. Staff from the DOE labs conducted systematic interviews about current practices in capacity expansion modeling, production-cost modeling, hydrological modeling, and transmission stability modeling while also discussing how scenario analysis is conducted and how models and data are integrated. The effort resulted in a set of model, integration, and scenario recommendations that could be valuable to TVA, other utilities, system operators, and other stakeholders conducting integrated grid analysis. Individual model recommendations suggest exploring computational tradeoffs with detail and resolution across spatiotemporal structure, supply- and demand-side details, transmission overlays, market interactions, and ancillary services. Automated processes to pass data between models and conduct larger scenario suites could also enhance valuation practices by enabling a more consistent study of asset value across a broader range of uncertain future grid conditions where PSH could be particularly valuable. TVA and other industry stakeholders can learn from and adapt applied research-grade methods developed by DOE laboratories and other research institutions to improve decision making and accelerate progress towards a reliable, economic, sustainable energy system.

13 HYDRO ENERGY↗

Rare Lepton Decays and Differentiable Hadronization Models - From Signatures of New Physics to Data-driven Event Generation

This dissertation is partitioned into two parts: phenomenological studies focused on rare lepton decays as probes of heavy and light new physics, and the development of differentiable, data-driven hadronization models. Part I develops the phenomenology of new physics signatures stemming from rare charged lepton flavor violating decays probed by experiments at the intensity frontier. These include interactions mediated by both high-scale effective operators and light new physics, manifesting in multi-lepton final states ($\mu \to 5e$), elastic nuclear transitions ($\mu \to e$ conversion), baryon-number-violating muon capture, and time-dependent signals from ultralight dark matter ($\mu \to e \phi, \tau \to \ell \phi$). Part II develops two distinct strategies for advancing differentiable and data-driven hadronization models. One involves comprehensive reweighting frameworks for hadronization that enable efficient uncertainty estimation, facilitate parameter tuning, and interface naturally with differentiable programming paradigms. The other introduces machine-learning-based methods for extracting microscopic fragmentation dynamics directly from macroscopic observables through the deformation of existing models -- effectively providing solutions to the inverse problem of hadronization. Altogether, these studies advance the interpretability, flexibility, and precision of theoretical predictions for both high-intensity and high-energy experiments.

Menzo, Tony [Cincinnati U.] (ORCID:000000022013457↗

Towards provision of regularly updated climate data from the Coupled Model Intercomparison Project

The Coupled Model Intercomparison Project (CMIP) is a flagship of the World Climate Research Programme (WCRP). CMIP has become a recognised ‘brand’ in climate circles evolving over the last thirty years from a targeted research activity by a small number of climate modelling centres intercomparing their Earth System Model (ESM) simulations to a broad international coordinated research effort (Durack et al, 2025). CMIP is organized as a research activity leveraging funded and in-kind contributions from experts within modelling centres and the broader scientific community supported more recently by a fully-funded International Project Office. Within CMIP, Model Intercomparison Projects (MIPs) are community-designed to understand past, present and future climate. CMIP data provides a valuable resource for climate research and is routinely used to assess model representation of climate processes and test scientific hypotheses in the context of model uncertainty and (forced and internal) variability as evident from its prolific use in scientific publications1 . The impact relies on enabling infrastructure (most prominently via the Earth System Grid Federation (ESGF)), which allows sharing of simulation output, provision of the boundary conditions used in each simulation, and definition of the data standards that are essential to facilitating wide use of the data. The impact is supplemented by the wide-ranging scrutiny to which model simulations are subjected. Beyond its use in research, CMIP data is a key resource for communities producing derived climate information from downscaling and impact studies, such as the Coordinated Regional Downscaling Experiment (CORDEX; Gutowski et al., 2016) and the Intersectoral Impacts MIP (ISIMIP; Frieler et al., 2024). Government, academic and commercial entities also increasingly rely on CMIP and its downstream data for climate risk assessments and climate services (for example, Copernicus Climate Change Service and World Bank portal). This means that, although CMIP is a research activity, it increasingly serves a secondary and very relevant role as a provider of climate data – a long-recognised dichotomy (Stevens, 2024). Research and applications have distinct needs, with the former requiring flexibility and generality and the latter consistency. Here we explain how the design of the research activity has been adapted to reduce the burdens imposed by applications and how the research infrastructure might evolve to further enable scientific inquiry. We propose one possible approach to consistently providing model information and projections for applications in the future.

Environmental sciences↗

Use of Rig Parameter Data in Bit Constraint Models for Improved Drilling Performance at The Geysers

Surface parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. However, these measurements are of reduced value without a standard to aid in evaluation and decision making. A method is demonstrated whereby drill bit constraint models are used to interpret drilling response parameters. Drill rig parameter data for well GDC-36 at the Geysers Geothermal Field Power were acquired by Geysers Power Company and drilling contractor Kenai Drilling using Pason US DataHub and evaluated. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) along with other model constraints in computational algorithms. The method is used to evaluate overall bit performance, monitor bit integrity, and detect the presence of drillstring vibrations and other conditions contributing to bit failure; comparisons are made to observations of bit wear and damage. The method will be applied in real-time to improve decision-making on subsequent wells and has applicability to development of advanced analytics on future geothermal wells using real-time electronic drilling recorder (EDR) data for improved performance and reduced drilling costs.

15 GEOTHERMAL ENERGY↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Adsorption of terbium (III) on DGA and LN resins: Thermodynamics, isotherms, and kinetics

Two commercially available extraction chromatography (EXC) resins containing N,N,N’,N’-tetra-n-octyldiglycolamide (DGA Resin, Normal, 50 – 100 μm) and Bis(2-ethylhexyl) phosphate (LN Resin, 100 – 150 μm) were used as adsorbents to study fundamental adsorption properties such as thermodynamic values, equilibrium isotherms, and kinetic uptake models for terbium(III) adsorption. Weight distribution ratios (D w ) for terbium on DGA and LN resins were measured using a [ 160 Tb]Tb 3+ radiometric tracer in nitric acid as a function of acidity, temperature, initial analyte concentration, and equilibrium time. The D w values showed increasing binding affinity for DGA resin at high nitric acid concentrations and decreasing binding affinity for LN resins. Thermodynamic studies for DGA and LN resins revealed that the Gibbs free energy (ΔG) increased consistently with temperature. To model equilibrium data, increasingly higher parameter equilibrium isotherm models (Henry (1) < Langmuir, Freundlich (2) < Redlich-Peterson (3) < Fritz-Schluender (4)) were compared on their root mean squared errors (RMSE) and adjusted determination coefficients to determine the most applicable model. In all cases, the empirical four-parameter Fritz-Schluender isotherm demonstrated a superior fit. Similar comparisons for reaction-based kinetic models (Pseudo-first-order < Pseudo-second-order < Pseudo-n-order) revealed that the higher-order PNO model yielded a superior fit of kinetic data for both resins. Furthermore, in some cases, adsorption isotherms and kinetic models could also be modeled by a lower-order model with minimal change in error parameters. Weber-Morris plots revealed that two linear sections are observed for each resin, where the first linear segment is attributed to fast (film diffusion) adsorption of terbium, followed by slower intraparticle diffusion of terbium through the pores as the rate-limiting step. Based on the Weber-Morris plot, both film and intraparticle diffusion are involved in controlling the kinetic rate of adsorption for DGA and LN resins.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Sensitivity Analysis of Drivers Water Shortage in the Los Angeles Region During Drought

The code and detailed step-by-step instructions for generating the model output data, processing results, and analysis and plotting are provided at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. The PyArtes model is a python adaptation of the Artes model. PyArtes uses many of the same input data and optimization model architecture as Artes. Documentation for the PyArtes model is provided in the Supplement to the paper. The primary data product are simulated monthly water shortages for indoor and outdoor demand under a large ensemble of drought scenarios (>13,000). The droughts are hypothetical and are not based on historical time series data of supply sources - though historical data did help inform ranges explored for supply parameters. Demands are informed by recent 2017-2021 water supply data. Demands used for the model can be accessed at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. Simulations resolve demand for over 90 water providers in the study region. The results report 36 months of water shortage data for each indoor and outdoor demand node. The study also developed a multilayer perceptron (MLP) neural network trained on a subset of the simulated shortage ensemble to emulate worst annual water shortage for a given set of parameter multipliers -- provided the parameter values fall within the ranges sampled in the ensemble. Emulated water shortages for synthetic ensembles are in the MLP-generated shortages folder. The MLP model was used to generate larger ensembles to support Sobol analysis that would have been extremely computationally expensive to simulate. Datasets provided in this repository*: Simulated shortages. These results are used for the analysis for Figures 5, 8, and 9 in the paper, and also to train the MLP emulator. .zip file containing outputs for the 13,312 scenario ensemble. Separate .csv files for indoor and outdoor shortage for each scenario. Rows = demand ids (~100), Columns = months (36) Units = acre-feet/month of shortage (shortage = monthly demand - supply). 1 acft = 1233.48 m^3 .csv files of aggregated shortages derived from the 13,312 ensemble Rows = scenarios (13,312), Columns = demand ids (~100) Units = acre-feet/year (either worst annual shortage or total shortage over the 3-year drought) .csv file of the parameter multipliers scenarios for the ensemble .csv file of the parameter ranges and baseline values the multipliers were applied to MLP-generated shortages. These results are used for Figures 4, 6, and 7 in the paper. mwd higher folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results Emulated shortages. Rows = scenarios, columns = demand ids, units acft Sobol results. Rows = demand ids, columns Sobol (S1, ST, or 95% confidence interval) value for each parameter mwd lower folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results same organization as mwd higher MLP performance: performance metrics (R^2, RMSE, BIAS, MAPE) for the testing subset (20% or 2,662 scenarios) and simulated vs emulated worst year shortage (acre-feet/year) for every demand node, MWD wholesale regions, and the entire study region (LAC). Supporting data for figures. Figure plotting scripts in the associated GitHub repo. These files support analysis and visualization. Geospatial Data used for plotting simulated water shortages and Sobol results. Dictionary of full names for demand nodes in the model and estimates of water supply by source type informed by Artes input files and California Urban Water Management Planning data: https://water.ca.gov/Programs/Water-Use-And-Efficiency/Urban-Water-Use-Efficiency/Urban-Water-Management-Plans *Readme files provided for each folder.

drought↗

National Energy Water Treatment & Speciation (NEWTS): A Water & Critical Mineral Database and Dashboard

The scarcity of water resources, the need for beneficial water reuse, and the challenges of wastewater treatment are becoming increasingly pressing in economic, social, and environmental domains. Addressing these concerns requires effective treatment strategies to manage wastewater streams and tackle environmental and economic issues. Furthermore, the recovery of critical minerals from the waste streams associated with energy production holds the promise of offsetting treatment costs and securing local sources of valuable minerals. However, relevant data on these waste streams are dispersed and challenging to locate. The process of ingesting such data into modeling software often involves multiple steps, requiring data restructuring to meet software-input requirements. The non-standardized reporting of water data makes data aggregation and reformatting a time-consuming process. Additionally, essential attributes necessary for modeling water treatment and mineral scale formation are frequently missing. Moreover, data gaps vary depending on the region of interest. Consequently, there is a pressing need for high-quality energy-water composition data that can be easily imported into water chemistry modeling software. To address this need, the National Energy Technology Laboratory has created the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard—a free online tool catering to community leaders and water researchers. NEWTS facilitates a comprehensive understanding of the composition of energy-related wastewater streams in the United States. The datasets provide detailed concentrations and speciation of major and minor aqueous compounds in energy-related wastewater streams, including power plant leachate, acid mine drainage, brackish water, and oil and gas produced water across the United States. Many of the aqueous species are critical minerals (Li, REEs) in high demand to modernize the world’s energy infrastructure. Many of the datasets also contain volumetric flow-rates needed to model the treatment and reuse scenarios in advanced aqueous chemistry software programs. The NEWTS Database and Dashboard offer public access to hitherto challenging-to-access datasets, presented in a standardized format that is tailored for easy input into aqueous chemistry modeling software. By performing the work needed to transform dispersed, disparate data sources into unified, model-ready datasets, NEWTS serves as an essential resource in advancing water treatment research and sustainable water resource management.

produced water management↗