Search NASASearch

SEARCH · Search NASA

Results for “loading models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Sensitivity Analysis of Numerical Modeling Input Parameters on Wind Turbine Loads in Deterministic Transient Load Cases

Aero-hydro-elastic-servo numerical models used to design and analyze wind turbines are based on thousands of variable input parameters that dictate the inflow, aerodynamic, structural, and control characteristics of the system as well as sea state, hydrodynamic, and mooring characteristics for fixed-bottom and floating offshore wind turbines. Each of these parameters has some level of uncertainty, which can significantly impact the predicted loads. Understanding the uncertainty in the inputs is critical to understanding the uncertainty in the outputs. This work demonstrates a screening technique to identify which parameters ultimate loads are most sensitive to so that more focus can be given to quantifying the possible range of those parameters. This technique has been demonstrated previously for different turbine and load case types and is extended here for a floating offshore wind turbine in design load cases with transient events both in the inflow and operations. Each load case features a deterministic gust, including variations in wind speed, direction, and shear. Load cases are considered with an operating turbine as well as with prescribed fault, startup, and shutdown procedures. The study found that key input parameters with a large impact on loads include the length of the gust, the magnitude of direction change and speed in the gust, the initial wind speed, and the shape of the gust profile.

17 WIND ENERGY

Dynamic Modeling and Simulation of a Subcritical Coal-Fired Power Plant under Load-Following Conditions

Dynamic models for power plants that capture realistic general process trends and effects of manipulated variables are needed to improve load-following, while minimizing carbon footprint. In this work, a dynamic modeling approach and simulation results for subcritical coal-fired power plant components are presented. These encompass simulation of the dynamics in the fireside, including the effects of fuel, air combustion, and the dynamics of the entire waterside and power generation sections. This model development enables the simulation and analysis of the important short and long-time scale dynamics of components such as heaters, evaporative loop, and power generation units. Furthermore, additional variables in the power generation section are introduced to improve model accuracy, extending the prediction capability of subcritical power plant models and opening new opportunities for research in operator training, optimization, and advanced model-based controller design that are based on these models. The change in process gain for different ramp rates associated with disturbance signals that affect process variables is also explored and a correlation developed. This provides opportunities to study disturbance rejection control implementation and adaptation for scenarios with such variations in ramp rates. The prediction capabilities of selected components are compared to data available in literature, with the obtained root mean squared error ranges that reflect the model performance and quality of predictions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Integrating Energy-Efficient Computing with Computational Research to Accelerate Energy Technology

NREL's computational sciences center hosts the largest high performance computing (HPC) capabilities dedicated to energy research while functioning as a living laboratory for energy-efficient computing. NREL's HPC capabilities support the research needs of the Department of Energy's Office of Energy Efficiency and Renewable Energy (EERE). In ten years of operation, HPC use in EERE-sponsored research has grown by a factor of 30, including work in electricity generation, energy efficiency, transportation, and energy system modeling. This paper analyzes this research portfolio, providing examples of individual use cases. The paper documents NREL's history of operating one of the world's most energy-efficient data centers while examining pathways to reduce economic and environmental impact beyond reduction of Power Usage Efficiency (PUE). This paper concludes by examining the unique opportunities created for accelerating improvements in data center efficiency created by combining an HPC system dedicated to energy research and a research program in energy-efficient computing.

97 MATHEMATICS AND COMPUTING

Data-Driven Modeling of High-Resolution Residential Load Profiles Using Low-Resolution Smart Meter Measurements

Accurate and high-resolution residential load profiles are essential for power system modeling, demand response planning, and effective grid operation. As the energy sector moves towards a more actively managed distribution system, the ability to understand residential energy consumption at a minute-by-minute scale becomes increasingly critical. High-resolution load profiles provide key insights into demand patterns and user behavior, enabling grid operators to design more effective energy solutions; however, residential load measurements in the field are typically recorded at low resolutions, such as 15-60 minutes, which makes it hard to study the characteristics of different residential customers. This paper addresses these challenges by introducing a data-driven approach to generate realistic, high-resolution residential load profiles based on lowre-solution measurements and weather information. The proposed method retains the key features of the actual residential load measurements while offering appliance-level energy consumption details for each residential building. The results demonstrate the effectiveness of the proposed load profile generator, proving its capability to support utilities in optimizing residential energy management and ensuring a more reliable and resilient grid.

24 POWER TRANSMISSION AND DISTRIBUTION

Thermo-hydraulic steam pipe models for district heating simulations: Simplifications to balance accuracy and simulation speed

Steam piping networks are essential for optimizing performance in industrial processes and district heating systems. However, dynamic models that balance thermo-hydraulic accuracy with computational efficiency remain limited. In response, this paper presents a new discretized steam pipe model based on the plug flow approach, capturing key thermo-hydraulic behaviors while simplifying steam phase change processes. Implemented in Modelica, the model accurately calculates temperature and pressure distributions along steam pipelines. To improve computational efficiency for district-scale simulations, five model simplifications are introduced: lumped thermo-hydraulic functions, empirical correlations, fluid state approximations, steady-state dynamics and inclusion of flow derivatives. These simplified models achieve 85%-98% accuracy in predicting pressure drop and condensation losses, including dynamic condensate behavior during pipe warm-up—a factor often overlooked in existing models. The models support diverse network configurations, scaling effectively to systems with multiple distribution pipes and connected building loads. Discrete models provide detailed insights but exhibit a cubic increase in simulation time as the network scales by N connected building O(N 2.42 ). In contrast, lumped models simulate 10–28 times faster than discrete, offering quadratic scaling of simulation time O(N 1.73 ). However, they still require 6 times more computation time than a lossless network, highlighting the inherent computational challenges of modeling compressible fluid flow. In conclusion, the steady-state lumped variant, with its near-linear scalability in computational time O(N 1.01 ), emerges as an efficient solution for preliminary design evaluations and extensive parametric studies.

15 GEOTHERMAL ENERGY

A generalizable machine learning-assisted fast Fourier transform algorithm to simulate the large strain phenomena in polycrystalline materials

Machine learning methods have shown initial promise in constitutive modeling for single crystals or homogenized polycrystals, delivering notable computational efficiency. However, existing machine learning-based constitutive models often lack generalizability, limiting their application across diverse boundary value problems. This study introduces a thermodynamics-informed artificial neural network model to accelerate rate-tangent crystal plasticity fast Fourier transform simulations for cross-scale deformation behaviors of polycrystals under complex loading. Our model integrates microstructural variability and local interactions effectively. To address local effects in each grain, we employ K-means clustering to group Gauss points within the microstructure into clusters assumed to be in similar mechanical states. This approach, based on self-clustering analysis, extends model scope from macroscopic stress response to the granular level, capturing mechanical responses and orientation evolution across grains. This reduces the number of nonlinear problems to solve, with cluster responses propagated throughout each group. The thermodynamics-based artificial neural network-extracted features are further processed using local material state clusters to account for history-dependent deformation and evolving microstructures. Additionally, representative volume element simulations with rate-tangent crystal plasticity fast Fourier transform provide reliable datasets for model training. The proposed model demonstrates high efficiency, accuracy, self-consistency, and enhanced generalizability in predicting strain–stress responses and orientation evolution at both individual grain and aggregate scales under complex loading conditions, such as biaxial tension and arbitrary loading scenarios.

36 MATERIALS SCIENCE

ACT University Collaboration Proposal: Role of manufacturing defects on material failure under dynamic loading for developing enhanced failure models and theories (Final Report)

Research on the tensile behavior of additively manufactured 316L stainless steel coupon specimens at increasing strain rates was conducted over the past year at Penn State. Dynamic loading rates into the 1000 1/s loading rates were performed on specimens with nearly 1.32” gage lengths. Stress-strain plots show ductile and plastic behavior well beyond 12% strain with necked specimens, having smaller effective gage lengths showing up to 56% ultimate strain. Additional tests performed on compact tension specimens helped with simulation work to understand the deformation behavior. Using finite elements, it was possible to determine the feasibility of a comprehensive experimental validation study towards an improved model for AM material failure – specifically, the Bai-Wierzbicki approach which accounts for lode angle and triaxiality. A non-significant number of tests are projected for eight specific failure nodes, with additional replicates to provide strain-rate capabilities to the existing formulation. One such correction factor is explored for quasi-static, notched specimens. Finally, stress intensity factor was explored using a set of compact-tension specimens which would also provide useful validation data for any simulation work.

36 MATERIALS SCIENCE

Coupled Aero-Hydro-Mechanical Hybrid Simulation Testing of Offshore Wind Turbines Subjected to Operational and Extreme Loading Conditions

Understanding the response of the Offshore Wind Turbine (OWT) subjected to realistic applied loads requires modeling the whole structure including its soil-foundation system. This requires unique and innovative testing facilities. OWT systems experience cyclic and dynamic loading due to wind, wave, current, rotor vibrations (i.e., 1P load) and vibrations caused by the blade shadowing effects (2P/3P loads). These loads are complicated in nature and have varying amplitudes, frequencies, and directions. Investigating the response of the entire OWT system including the soil-foundation system under these complex loading conditions, requires: (1) full understanding of the loading characteristics including: the power take-off mechanical load (1P and 3P), and areo- and hydrodynamic loads that the OWT system is subjected to; (2) testing facility with unique multidirectional loading capabilities that allows for simultaneous application of realistic wind, wave and machine loads, axial gravity loads, and induced overturning moments; and (3) unique and cost-effective testing techniques that allow for accurate analysis of the overall response of the OWT system under realistic conditions such as: Real-Time Hybrid Simulation (RTHS).

17 WIND ENERGY

popclass: A Python Package for Classifying Microlensing Events

popclass is a Python package that provides a flexible, probabilistic framework for classifying the lens of a gravitational microlensing event. Gravitational microlensing occurs when a massive foreground object (e.g., a star, white dwarf or black hole) passes in front of and lenses the light from a distant background source. This causes an apparent brightening, and shift in position, of the background source. In most cases, characteristics of the microlensing signal do not contain enough information to definitively identify the lens type. Different lens types lie in different but overlapping regions of the characteristics of the microlensing signal. For example, black holes tend to be more massive than stars and therefore cause microlensing signals that are longer. Current Galactic simulations enable us to predict where different lens types lie in the observational space and can therefore be used to classify events (e.g., Lam et al., 2020). popclass allows the user to match the characteristics of a microlensing signal with a simulation of the Galaxy to calculate lens type probabilities for the event (see Figure 1). Constraints on any microlensing signal properties and any Galactic model can be used. popclass comes with an interface to ArviZ (Kumar et al., 2019) and PyMultiNest (Buchner et al., 2014) for microlensing signal constraints, as well as pre-loaded Galactic models, plotting functionality, and methods to quantify the classification uncertainty. The probabilistic framework for popclass was developed in Perkins et al. (2024), used in Fardeen et al. (2024) and has been applied to classifying events in Kaczmarek et al. (2025).

97 MATHEMATICS AND COMPUTING

California Price Response Potential Study

California's energy landscape is undergoing a significant transformation, driven by the increasing integration of renewable energy sources, the increased adoption of distributed energy resources, the electrification of end-use loads, and the growing need for grid efficiency. To address these challenges, recent revisions to the State’s Load Management Standards (LMS) require all of California’s large utilities and community choice aggregators (CCAs) to offer dynamic electricity pricing options to customers by 2027. Dynamic pricing, which involves varying electricity rates based on real-time supply and demand conditions, offers a promising solution for optimizing grid operations, reducing costs, and incentivizing efficient use of grid capacity. Effective implementation of dynamic pricing requires understanding the potential impacts on customer bills, system load, and the cost-effectiveness of automation technologies. This study aims to evaluate the load response of various end-use devices to hourly dynamic prices. The end-uses studied here are space cooling, space heating, water heating, crop irrigation, pool and spa pumps, and electric vehicle (EV) charging, all for both residential and commercial applications, except for crop irrigation. In 2030, these end uses are forecasted to account for 18% of annual electricity demand in the state, but 40% of demand in the peak net load hour. By modeling possible price-responsive load dispatch algorithms and assessing the resulting impacts on both individual bills and the overall grid, we seek to inform policymakers and utilities about the potential benefits and challenges associated with dynamic pricing, and considerations for the design of dynamic pricing tariffs. Additionally, we will explore the cost effectiveness of adopting automation technologies to enable devices to respond more effectively to real-time price signals. This study considers a range of price profiles, accounting for differences across utilities and customer classes, and presents scenarios for dynamic price design via variation in the percentage of total customer electric costs that are allocated dynamically (versus constituting a fixed portion of the hourly volumetric price). We present results focused primarily on 2030, forecasting electricity prices under both low and high-cost scenarios, to inform longer-term tariff design considerations. We design tariffs by starting with 2019 prices that were calculated according to CalFUSE guidance (CPUC, 2022) and that have been used in recent studies; these prices are all-in volumetric rates that vary by utility and are revenue-neutral to each customer class. They are developed by considering six electricity cost components that are allocated hourly based on system load indicators (gross and net load, and wholesale prices). These prices are forecasted to 2030 for low and high cost scenarios, considering recent trends in total electricity costs with and without years of substantial wildfire mitigation investments. These tariffs, which allocate all costs on an hourly basis, are considered our “Full” dynamic tariff design scenario, while two additional scenarios explore allocating a portion of costs as a flat volumetric charge: the “Medium” scenario allocates 50% of revenue dynamically (and keeps 50% flat), while the “Mild” scenario allocates 20% of revenue dynamically. The 20% dynamic allocation on the Mild scenario aims to represent a case where only the marginal operating costs of the grid are included in the dynamic price.

29 ENERGY PLANNING, POLICY, AND ECONOMY

IDAES-PSE 2.7.0 Release

The Institute for the Design of Advanced Energy Systems (IDAES) Integrated Platform is a versatile computational environment offering extensive process systems engineering (PSE) capabilities for optimizing the design and operation of complex, interacting technologies and systems. IDAES enables users to efficiently search vast, complex design spaces to discover the lowest cost solutions while supporting the full process modeling lifecycle, from conceptual design to dynamic optimization and control. The extensible, open platform empowers users to create models of novel processes and rapidly develop custom analyses, workflows, and end-user applications. IDAES-PSE 2.7.0 Release Highlights New features: AutoScaler and CustomScalerBase classes: Such tools are the core of the new scaling framework being implemented in IDAES. Wider adoption of scaling tools among users will result in quicker and more robust model solutions. Scaler for equilibrium reactor and saponification properties: These scaler models are examples to follow for how to use the new scaling tools. ONNX Surrogate support from Optimization & Machine Learning Toolkit (OMLT): ONNX is an open standard format to save and load ML/AI models that is widely supported by all major frameworks. This capability makes it easier for IDAES users to create surrogate models and use them without having to support each framework individually. 1D Membrane Model for CO2 Capture and Utilization: Supports ongoing efforts for modeling and optimizing polymer membrane processes for CO2 capture and conversion into formic acid. StreamScaler unit model: Unrelated to the CustomScalerBase, this unit model allows a stream’s extensive variables to be scaled by a fixed factor. This allows streams being processed by multiple units in parallel to be scaled down to unit scale and scaled back up to process scale. Bug fixes or improvements: Scaling, EoS, Diagnostics tool, Modular Properties, tests & documentation Deprecations: Old Cubic EoS

AS

THERMAL MODELING OF HANFORD CESIUM AND STRONTIUM CANISTERS DURING SIMULATED LOADING

A computational fluid dynamics (CFD) model was built to simulate planned testing of heater assemblies within a canister and overpack for the Hanford Lead Canister (HLC) project. The HLC is a canister storage system that will contain heaters to simulate the decay heat of nuclear material and provide the canister storage system with environmental conditions equivalent to the operating conditions on a dry storage pad. The HLC will be equipped with long-term data collection and monitoring systems to provide an early warning of corrosion, pitting, cracking, or other signs of canister degradation that might threaten the integrity of the containment boundary over the potentially long term of dry storage. An important part of the HLC development is to make pretest numerical predictions for the behavior of the heated canister during the simulated radiolytic decay heat testing, which simulates the dry storage system during loading operations. The simulated radiolytic decay heat test is planned for mid-2024 in a configuration that includes the heater assembly, overpack, and canister, but with the lids removed to allow loading cesium and strontium capsules into the canister. One of the goals of the test is to evaluate the thermal behavior of the canister and overpack assembly in the ambient air of the test facility, which will provide data critical to validating the thermal models and understanding how the HLC will perform as a system once deployed. To best approximate real-world conditions, the CFD model includes the full air volume of the mock-up truck bay the heated canister test will be performed in, enabling detailed investigation of how the heated canister affects airflow around it. Rigorous pre-deployment testing of the complete HLC cask and canister system is intended to be completed before the HLC is deployed in the 2028 timeframe. This study presents the pre-test temperature predictions of the simulated radiolytic decay heat test. A description of the heater assembly, canister, and overpack system is presented. The model was developed with the commercial CFD software STAR-CCM+. An uncertainty analysis was run with the CFD model to determine the uncertainty in the temperature predictions and provide a range over which the predicted temperatures are expected to vary. The uncertainty analysis was preformed by coupling STAR-CCM+ with the software Dakota, which provides advanced parametric analyses, including quantification of margins and uncertainty with computational models. This work is expected to provide insight into SNF canister behavior.

Carpenter-Graffy, Dina E.

Scalable Risk Assessment of Rare Events in Power Systems With Uncertain Wind Generation and Loads

Risk assessment of rare events has become increasingly important in power system planning and operation with the increasing integration of renewable energy and the presence of system uncertainties. However, quantifying the risk posed by rare events via the traditional method, i.e., Monte Carlo sampling (MCS), incurs substantial computational expense stemming from the vast ensemble of power flow simulations. To accelerate the assessment, this paper proposes a Deep Neural Network (DNN)-kernelized vector-valued Gaussian Process (VVGP) approach with excellent computational efficiency while maintaining high accuracy. Consequently, serving as a surrogate model for the power flow solver, the DNN-kernelized VVGP enables significantly faster but accurate risk assessment compared to the power flow solver. The developed surrogate model evaluates low-order N - k events that contain more than 90% instances by adeptly capturing the topological features while the high-order N - k events are assessed via a power flow solver, thereby striking a balance between computational efficiency and uncertainty quantification accuracy. Moreover, the model incorporates a Support Vector Machine (SVM) classifier to resample concerning low-probability tail events to counteract the biases potentially introduced during the DNN-kernelized VVGP evaluations. Simulations conducted on the modified IEEE 24-bus, 118-bus, and European 1354-bus systems demonstrate that the proposed method maintains the accuracy benchmark set by MCS while significantly reducing computational demands in large-scale power systems as compared to other state-of-the-art methods.

17 WIND ENERGY

Integration of Electric Vehicle Charging Loads in Residential Building Stock Energy Modeling

The rapid adoption of electric vehicles (EVs) has resulted in significant new household electric loads that have the potential to change how energy costs are incurred by homeowners and the landscape of utility operations and energy infrastructure. Whereas adoption patterns and magnitudes of residential building and EV charging loads are influenced by distinct factors, the loads themselves are tightly coupled with the behavior of the individual occupants and EV owners.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

IM3 Data Center Driven Grid Stress Dataset for the U.S. Western Interconnection

This dataset provides projected grid stress and reliability results (including all model inputs and outputs from an open-source grid operations modeling framework - GO), for the Integrated Multisector, Multiscale Modeling (IM3) project, under varying levels of data center demand growth between 2025 and 2035 in the U.S. Western Interconnection. The scenarios and sensitivity experiments are combinations of different data center demand growth rates and energy, weather, population and economic pathways. Data center demand growth projections were sourced from the Electric Power Research Institute (EPRI). The data center demand growth projection names are: Low (3.71% annual data center demand growth) Moderate (5% annual data center demand growth) High (10% annual data center demand growth) Higher (15% annual data center demand growth) Energy, weather, population and economic pathways are informed by two Shared Socioeconomic Pathways (SSP3 and SSP5) and two Representative Concentration Pathways (RCP4.5 and RCP8.5) following the hotter general circulation model (GCM) forcing group from a set of perturbed thermodynamics simulations. The resulting pathway names are: rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 The main scenarios and sensitivity experiments are detailed below. Reference scenario: The projected grid stress and reliability results for the U.S. Western Interconnection from a previous study. This scenario does not consider data center demand growth explicitly. Data center scenario: Building on the reference scenario, this scenario considers various data center growth rates and how they impact the U.S. Western Interconnection. Data center loads are modeled as flat 8760-hr profiles. This scenario does not consider new generation and transmission capacities specifically designed to meet the new data center demands. The related folder is named "flat". Delayed generator retirements sensitivity experiment: Building on the data center scenario, this experiment explores the impact of different levels of natural gas and nuclear generator retirement delays. The resulting scenario names are: (1) postponing 100% nuclear retirements; (2) postponing 100% nuclear and 25% natural gas retirements; (3) postponing only 50% natural gas retirements; (4) postponing 100% nuclear and 50% natural gas retirements; (5) postponing 100% nuclear and 75% natural gas retirements; and (6) postponing 100% nuclear and 100% natural gas retirements. The related folder names are: no_gen_retire_0_gas, no_gen_retire_25_gas, no_gen_retire_50_gas, no_gen_retire_50_gas_only, no_gen_retire_75_gas, and no_gen_retire_100_gas. Demand response through curtailment sensitivity experiment: Building on the data center scenario, this experiment explores the impact of different participation and compensation levels of data center demand response. The resulting scenario names are: (1) 5% demand available for curtailment with 750 $/MWh compensation; (2) 5% demand available for curtailment with 500 $/MWh compensation; (3) 5% demand available for curtailment with 250 $/MWh compensation; (4) 15% demand available for curtailment with 750 $/MWh compensation; (5) 15% demand available for curtailment with 500 $/MWh compensation; and (6) 15% demand available for curtailment with 250 $/MWh compensation. The related folder names are: dr_cost_250_drup_0_drdown_5, dr_cost_250_drup_0_drdown_15, dr_cost_500_drup_0_drdown_5, dr_cost_500_drup_0_drdown_15, dr_cost_750_drup_0_drdown_5, and dr_cost_750_drup_0_drdown_15. Combination of delayed generator retirements and demand response through curtailment sensitivity experiment: The impact of combining postponing 100% nuclear and 25% natural gas retirements with 5% demand available for curtailment with 750 $/MWh compensation is simulated. The related folder is named "dr_cost_750_drup_0_drdown_5_nuc_100_gas_25". Please refer to the README file for a detailed description of the dataset including individual files and references.

Artificial Intelligence

Understanding Electric Vehicle Range and Charging Needs: Interactions Between Ambient Temperature, Commute Patterns, and State-of-Charge Usage

Electric vehicle (EV) performance can vary substantially under real-world operating conditions, particularly due to ambient temperature effects on energy consumption, battery behavior, and thermal management requirements. This study quantifies how weather conditions, daily driving patterns, and State-of-Charge (SOC) usage strategies jointly influence EV driving range, charging frequency, and overall energy efficiency. A detailed and experimentally validated Autonomie vehicle model is developed, integrating a powertrain, a mono-zonal cabin model, and a battery electro-thermal model. Three battery sizes (200-, 300-, and 400-mile homologated ranges) are assessed across five commute profiles (20–200 miles) and six ambient temperatures (−18 °C to 50 °C), including scenarios with and without preconditioning. Results show that extreme temperatures could significantly decrease the maximum achievable range by up to 55% in cold conditions (−18 °C) and 40% in hot conditions (50 °C), relative to moderate conditions. Larger battery packs retain a greater fraction of their nominal range under thermal stress, while smaller packs experience sharper relative penalties due to the higher contribution of thermal loads to total energy demand. The analysis further demonstrates that limiting operation to partial SOC windows (e.g., 80–20%), a common real-world practice, significantly reduces achievable range and increases charging frequency, particularly in cold weather. Thermal preconditioning while plugged in is shown to mitigate these effects for short trips, reducing energy consumption by up to 31% in hot conditions and 7% in cold conditions. The findings demonstrate how climate, SOC usage behavior, and thermal management jointly shape the practical driving capability of EVs, highlighting the importance of efficient thermal management and realistic user charging strategies for ensuring reliable EV operation across diverse climatic scenarios.

33 ADVANCED PROPULSION SYSTEMS

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE