Search NASASearch

SEARCH · Search NASA

Results for “Ising model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,210 records · Page 5

Commutative Algebra Modeling in Materials Science – A Case Study on Metal–Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry’s constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

Khaemba, Caleb S.

Used Nuclear Fuel Management Using the Next Generation System Analysis Model

The U.S. Department of Energy (DOE) is leading the National effort to manage the back end of the nuclear fuel cycle, encompassing the safe transportation, storage/staging, and/or eventual disposal of used nuclear fuel (UNF) and high-level radioactive waste. The Next Generation System Analysis Model (NGSAM) is DOE’s discrete-event, agent-based simulation tool designed to model the full life cycle of UNF from reactor discharge to final disposal. NGSAM supports the DOE Office of Spent Fuel and High-Level Waste Disposition by enabling a detailed, scenario-based analysis of logistics, infrastructure, and shipping strategies. NGSAM replaces legacy models with a modern, flexible platform built on Repast Simphony and enhanced by the Process Analysis Tool. NGSAM simulates the movement and interaction of individual fuel assemblies with system components such as canisters, casks, railcars, and facilities. The model integrates with the Java Transportation Operations Model to plan and execute transportation scenarios, supporting both constrained and unconstrained resource allocation. Key features include customizable allocation and acceptance algorithms, detailed facility-level operations, and a Quick Edit tool for rapid scenario adjustments. NGSAM supports multimodal transportation modeling (e.g. rail, road, barge) and provides comprehensive cost, schedule, and infrastructure data. NGSAM utilizes data from sources such as DOE’s STANDARDS UNF database and DOE’s Stakeholder Tool for Assessing Radioactive Transportation, while also allowing user-defined inputs for scenario customization. NGSAM enables stakeholders to evaluate complex UNF management strategies, assess system performance under varying assumptions, and inform decision making for future infrastructure investments. Its modular architecture and integration with other Integrated Waste Management System tools make it a critical asset for planning the safe and efficient disposition of the Nation’s growing UNF inventory.

Craig, Brian [Argonne National Laboratory (ANL)]

Computational modeling of coupled mechanical damage and electrochemistry in ternary oxide composite electrodes

Performance degradation of ternary layered oxide cathodes largely originates from their loss of structural integrity in cyclic usage. Mechanical damage, such as intergranular fracture of the active particles, is not only a mechanical cleavage process but also interferes with electrochemical kinetics such as infiltration of liquid electrolyte, surface corrosion of the constituent primary particles, and may eventually isolate the primary grains from the electron conducting network. Here, in this work, we develop a computational framework that integrates electrochemistry of a LiNi x Mn y Co 1−x−y O 2 (NMC) composite cathode with mechanical damage of the active particles. To fully examine the intricate chemomechanical behavior of the electrode, we evaluate the effects of the anisotropic material properties, the influence of mechanical potential on Li transport, and the concurrent intergranular fracture and electrolyte penetration along the grain boundaries upon multiple cycles. Electrolyte infiltration benefits capacity retention but aggravates further mechanical damage by corrosion. Structural failure mostly occurs in the first charging due to the anisotropic mechanical strain between the primary grains, while the resulting damage remains stable in the later few cycles. The results are consistent with experimental observations and the integration of electrochemistry and mechanical failure enables a step further understanding of the complex mechanism of battery degradation.

Battery degradation

Spectroscopy in Nanoscopic Cavities: Models and Recent Experiments

The ability of nanophotonic cavities to confine and store light to nanoscale dimensions has important implications for enhancing molecular, excitonic, phononic, and plasmonic optical responses. Spectroscopic signatures of processes that are ordinarily exceedingly weak such as pure absorption and Raman scattering have been brought to the single-particle limit of detection, while new emergent polaritonic states of optical matter have been realized through coupling material and photonic cavity degrees of freedom across a wide range of experimentally accessible interaction strengths. In this review, we discuss both optical and electron beam spectroscopies of cavity-coupled material systems in weak, strong, and ultrastrong coupling regimes, providing a theoretical basis for understanding the physics inherent to each while highlighting recent experimental advances and exciting future directions.

Chemistry

Hierarchical Multi-agent Large Language Model Reasoning for Autonomous Heterogeneous Catalyst Discovery

Artificial intelligence is reshaping scientific exploration, but most methods automate procedural tasks without engaging in scientific reasoning, limiting autonomy in discovery. We demonstrate that hierarchical agentic large language model reasoning can efficiently drive simulation and scientific exploration. Across two chemical applications, CO adsorption on Cu surface transition metal adatoms and on M–N–C catalysts, reasoning-guided exploration reduces required atomistic simulations by up to 90% relative to heuristic or random selection. Comparisons across single-agent, multi-agent, and stochastic baselines show that hierarchical strategies yield more coherent and information-efficient search trajectories. Reasoning traces reveal chemically grounded decisions that cannot be explained by semantic bias or stochastic sampling. We realize these agentic reasoning strategies in Materials Agents for Simulation and Theory in Electronic-structure Reasoning (MASTER), a multimodal system that translates natural language into density functional theory workflows. Altogether, multi-agent collaboration accelerates heterogeneous catalyst discovery and marks a step toward more autonomous, reasoning-guided scientific exploration.

30 DIRECT ENERGY CONVERSION

Design and Modeling of the Charge Readout of a SiMOS Quantum Dot with a Single Electron Transistor and CryoCMOS

Single electron spin qubits trapped in SiMOS quantum dots are a promising technology for scaling to thousand- or million-qubit systems due to their compatibility with mature CMOS manufacturing processes. A readout system that combines a single electron transistor with a custom cryogenic CMOS amplification and digitization chain offers key advantages by avoiding the use of bulky RF components or room-temperature interconnects. We present design techniques and simulation results for an optimized qubit-SET cryoCMOS interface, culminating in the design of the QNDR1 ASIC, the first cryogenic readout ASIC designed under the Quandarum project, which targets the development of a many-channel spin qubit based detector for use in high energy physics.

Quinn, Adam [Fermilab] (ORCID:0009000605056371)

Advanced Model Development for Large Eddy Simulation of Oxy-Combustion and Supercritical Carbon Dioxide Power Cycles

A joint experimental and numerical study is performed to observe the characteristics of a supercritical carbon dioxide turbulent mixing layer in the presence of strong nonlinearities in the thermodynamic and transport properties. A bespoke experimental setup is designed and employed for this purpose and provides insight into macroscopic mixing behavior. The mixing is experimentally observed using two techniques: shadowgraphy and spontaneous Raman scattering. Qualitative and quantitative intensity fields obtained via these techniques yield instantaneous and mean density data. Spanwise temperature data is also collected using analogue resistance temperature detectors. These measurements are used to quantify the level of mixed material within the field. The experimental data are supplemented by a companion high-fidelity numerical study. The numerical results are obtained through fully resolved, three-dimensional direct numerical simulation. The numerical dataset permits observation of the near-field mixing characteristics, which are difficult to measure experimentally due to the rapid dynamics and sharp thermophysical gradients in this area. Qualitative field visualizations are presented, followed by quantitative mixed material results and observations regarding thermodynamic property trends at select locations within the field. One-dimensional spectra of the turbulent kinetic energy and solenoidal dissipation are provided to observe the spectral characteristics of the flow. Reynolds stress anisotropy is analyzed graphically through anisotropy invariance maps (Lumley triangles). The mixing quantification, spectral data and anisotropy analysis of a flow at these thermodynamic conditions represent the main outcomes of the work.

20 FOSSIL-FUELED POWER PLANTS

Precise Modeling of a Complex Solenoidal Magnetic Field Using a Combination of Analytic Functions and a PINN

We demonstrate an iterative approach to modeling a sparsely measured magnetic field in a large-bore solenoid. This approach uses a hybrid of traditional and machine learning techniques. The traditional technique is a linear least-squares fit using a series solution to Laplace's equation, while the machine learning technique involves the training of a physics-informed neural network (PINN) on the least-squares fit residuals. We use a newly defined activation function "DELTAsnake," a modification to the snake activation function proposed by Ziyin et al. that allows for stronger curvature and non-monotonicity. The combined model approximately obeys Maxwell's equations to a level sufficient for producing high quality physics simulations and analysis. Our approach is applied to a highly realistic calculation of the expected magnetic field in the Mu2e experiment's Detector Solenoid which includes a simple model for the expected statistical measurement uncertainties. Using ten toy measurement simulations, we demonstrate the capabilities of our model in comparison to the least-squares method alone; the least-squares method alone results in a reduced chi-squared statistic of ${2.15 \pm 0.01}$, while our approach improves the reduced chi-square to ${1.034 \pm 0.005}$. Furthermore, for an average toy simulation, we show that the range of the RMS of the three field component residuals reduces from ${0.07-0.37}$ Gauss to ${0.05-0.07}$ Gauss. We find that this novel method is robust against a realistic systematic uncertainty deriving from Hall probe calibration bias and can be used to significantly reduce the number of measurements required to achieve an accurate model.

Kampa, Cole [Caltech] (ORCID:0000000192972920)

Computing the Critical Temperature of the Affine-Transformed $D=3$ Ising Model Using Masked Autoregressive Flow

The simple Ising model provides a rich environment to build and study lattice field theories. As part of an ongoing project to construct a conformal field theory (CFT) on an arbitrarily curved manifold, in this work we develop methods to measure the critical temperature $β_c$ of the affine-transformed Ising model on the face-centered cubic (FCC) lattice. The main challenge in this endeavor is finding a computationally efficient and accurate method of interpolating and extrapolating Monte Carlo observables with respect to coupling coefficients and temperature. Herein, we compare two such methods. A traditional statistical approach uses the multiple histogram (MH) method, while a newer machine learning approach uses a masked autoregressive flow (MAF) to estimate the underlying probability density function of a set of observables. While the MH method is specifically designed to interpolate and extrapolate Monte Carlo observables, we find that MAF is a viable alternative for measuring $β_c$ with a computational cost that scales more favorably. Furthermore, we comment on additional advantages of MAF relevant to our work, such as extrapolating in system volume.

Svenson, Kai [Texas U.]

Modeling the Potential Impact of Storage on the US Power Sector in a Multisector Dynamic Context

The electric power sector is expected to grow in size, importance, and complexity around the world as economies expand and electric supply technologies and demand patterns evolve in significant ways. At the same time the electric power sector may see substantial increases in VRE generating technologies which can add variability and uncertainty to the diurnal and seasonal profile of electric supply. With these forces in play, the emergence of modular, flexible electricity storage technologies may have profound impacts on the operation of electric power systems. Here we reduce a storage modeling gap in MSD models by incorporating grid-based electricity storage into the electric sector dynamics of the Global Change Analysis Model-USA (GCAM-USA) with improved power sector representation. We find a potentially significant role for storage technologies in the future of the U.S. power system, with storage capacity ranging from 7.7 to 14.7 GW in 2050 and 13.9 to 31.6 GW in 2100 across several techno-economic scenarios. We also find that storage can help to smooth variability in residual load arising from evolving electricity demands and the introduction of high shares of VRE to the electric grid. This reduces reliance on high-cost peaking generators, improves system-wide capacity factors, and limits curtailment of VRE.

Patel, Pralit L.

A quality-agnostic combinatoric cost estimation model for large-format directed energy deposition metal additive manufacturing

Directed energy deposition (DED) additive manufacturing (AM) processes are amenable to synergistic combination into multi-process AM systems due to similar requirements for automation and energy sources. This work analyzes the economic performance of such DED AM systems from a quality-agnostic combinatoric standpoint with a model that calculates lowest-cost system combinations based on part geometry and process performance metrics. Common DED AM systems research focuses on a single process and does not consider the process, system, and application in the context of all possible system combinations (e.g., the combined set of process selection(s), motion system(s), and process hardware), leading to limited applicability of the resulting DED AM systems to cost-sensitive components such as those found in energy generation applications. The model developed herein incorporates the capital, material, and energy costs associated with DED AM system combinations into a predictive tool for estimating part and system cost, the output of which is intended to guide deployment of finite research and development resources towards DED AM system combinations with the lowest costs and greatest likelihood of economic impact. The DED AM systems identified by this framework may enable domestic production of the large conventionally cast and forged components necessary for energy generation.

Shanafield, Alexandra [ORNL]

Phase-field modeling of stored-energy-driven grain growth with intra-granular variation in dislocation density

Abstract We present a phase-field (PF) model to simulate the microstructure evolution occurring in polycrystalline materials with a variation in the intra-granular dislocation density. The model accounts for two mechanisms that lead to the grain boundary migration: the driving force due to capillarity and that due to the stored energy arising from a spatially varying dislocation density. In addition to the order parameters that distinguish regions occupied by different grains, we introduce dislocation density fields that describe spatial variation of the dislocation density. We assume that the dislocation density decays as a function of the distance the grain boundary has migrated. To demonstrate and parameterize the model, we simulate microstructure evolution in two dimensions, for which the initial microstructure is based on real-time experimental data. Additionally, we applied the model to study the effect of a cyclic heat treatment (CHT) on the microstructure evolution. Specifically, we simulated stored-energy-driven grain growth during three thermal cycles, as well as grain growth without stored energy that serves as a baseline for comparison. We showed that the microstructure evolution proceeded much faster when the stored energy was considered. A non-self-similar evolution was observed in this case, while a nearly self-similar evolution was found when the microstructure evolution is driven solely by capillarity. These results suggest a possible mechanism for the initiation of abnormal grain growth during CHT. Finally, we demonstrate an integrated experimental-computational workflow that utilizes the experimental measurements to inform the PF model and its parameterization, which provides a foundation for the development of future simulation tools capable of quantitative prediction of microstructure evolution during non-isothermal heat treatment.

Materials Science

Impact of Crystalline Phases on Low-Activity Waste Glass Durability: Insights from PCT and VHT

During vitrification of nuclear wastes, slow cooling along the container centerline promotes crystalline phase formation, which can alter residual glass composition and reduce chemical durability. This study investigates the effects of crystalline phases on the chemical durability of low-activity waste (LAW) borosilicate glasses using the product consistency test (PCT) and vapor hydration test (VHT) on container centerline cooled (CCC) samples. A preliminary model (R2 = 0.88) was developed to predict CCC PCT responses based on glass composition, PCT data from quenched glasses, and measured crystal fractions. Using the latest LAW glass dataset, the feasibility of predictive modeling is evaluated, limitations in current data and methods are identified, and challenges for improving model accuracy are discussed to guide future data collection and model development.

borosilicate glass

Energy-Optimal Vehicle Longitudinal Motion Control via Pontryagin’s Minimum Principle and Ultra-Local Model

Longitudinal vehicle motion control is essential for enhancing performance and optimizing a vehicle’s energy usage. However, it remains a challenging task due to the nonlinear and uncertain nature of vehicle dynamics, along with varying driving conditions. This paper presents a novel ultra-local optimal control approach based on Pontryagin’s Minimum Principle (PMP) that circumvents the need for detailed system identification by employing an ultra-local model. The control objective is to minimize the total energy consumption under boundary conditions while ensuring smooth traction force generation. The proposed approach is evaluated using a high-fidelity vehicle model in three representative scenarios: (i) nominal driving, (ii) a change in tire road friction coefficient (TRFC) from 0.5 to 0.65 and road slope from 0% to 5% during the maneuver, with target velocity unchanged, and (iii) a change in target velocity from 20 m/s to 0 m/s during the maneuver, while maintaining nominal TRFC and slope conditions. The simulation results demonstrate that the proposed method delivers robust performance, effectively balancing consumption and tracking accuracy in all tested scenarios.

Waleed khan, Muhammad [The University of Texas at

High-resolution modeling of indoor radon exposure with uncertainty quantification in Utah

Indoor radon accounts for 37% of population-level exposure to ionizing radiation in the United States. However, radon metrics are typically reported at coarse spatial scales, potentially obscuring meaningful local variation. We developed a high-resolution modeling framework to estimate indoor radon concentrations across Utah while explicitly quantifying predictive uncertainty. A total of 19,497 residential radon measurements collected between 2006 and 2017 were combined with environmental and housing characteristics and analyzed using a geospatial neural network that accommodates spatial dependence and nonlinear associations. Predictions were generated on a uniform hexagonal grid at 0.73 km2 resolution (H3 level 8). Out-of-sample predictions aggregated to the H3 level 8 grid showed good agreement with observed concentrations (Pearson r=0.64), while household-level predictions exhibited more moderate agreement (r=0.45). The model produced well-calibrated uncertainty estimates, with 24.1% of held-out observations exceeding the predicted 75th-percentile threshold. Maps of predicted radon concentrations and the probability of exceeding the U.S. EPA action level of 148 Bq/m3 (4 pCi/L) revealed substantial fine-scale spatial heterogeneity that was not apparent in conventional coarse-resolution summaries, with greater local variability observed in densely monitored urban counties than in sparsely sampled regions. High-resolution radon models that explicitly quantify uncertainty provide a useful framework for characterizing the spatial distribution of indoor radon and identifying areas of elevated exceedance risk. These findings highlight the value of fine-scale monitoring data and uncertainty-aware modeling approaches for radon exposure assessment, environmental risk characterization, and radon-related health research.

Wu, Yunhan [ORNL] (ORCID:0000000178842994)

Prime Time for Model-Predictive Control? Assessing the Technical and Market Readiness of Advanced Controls in Buildings

Despite three decades of extensive research and field testing that have consistently validated the benefits of Model Predictive Control (MPC) in building applications, the technology has seen limited market adoption. This paper evaluates the readiness of MPC for widespread deployment, showcases recent demonstrations and field tests across diverse building types, including residential, small commercial, large commercial, and campus settings. Our results demonstrate that MPC can optimize system operations to achieve load shifting, minimize curtailment of on-site generation, and reduce energy costs by up to 80 %, while maintaining or improving occupant comfort. We also show that MPC can effectively control large assets, such as MW-sized thermal storage systems, and respond to dynamic pricing signals. However, achieving scale remains difficult due to labor-intensive workflows, reliance on a “PhD-in-the-loop” for MPC design and maintenance, susceptibility to fragile data infrastructure, and persistent workforce education and acceptance barriers. To bridge this gap, we outline a transition from bespoke, labor intensive prototypes toward streamlined, segment-targeted deployment strategies that leverage model templates, semantic tools, and generative AI. By automating control configuration and reducing engineering effort, these recommendations provide a pathway for transforming successful research demonstrations into scalable, market ready solutions for MPC-based controls.

Pritoni, Marco

Uncertainty quantification for competing failure mechanisms in unidirectionally reinforced carbon–carbon composites

Microstructure-informed finite element models play a key role in the carbon–carbon composite design process. Variability in manufacturing process parameters and experimental limitations introduce model parameter uncertainty. This study quantifies the effect of model parameter uncertainty on transverse tensile fracture behavior and proposes a methodology to predict the failure mode based on competing microscale damage mechanisms. Finite element simulations incorporate fiber–matrix interface debonding with cohesive zones and matrix damage with a smeared crack band approach in a unidirectional carbon–carbon composite. Results from a variance-based global sensitivity analysis identifies interfacial and matrix damage parameters as the primary source of variability in fracture behavior. Sobol’ indices indicate that matrix and cohesive zone strengths contribute 94% of the variance in the effective ultimate stress. A local analysis elucidates the relationship between these constituent strength parameters and failure mode by estimating the probability of cohesive, matrix, and mixed-mode dominated failure. Based on the results for 4000 simulations, 93% exhibit mixed-mode or interfacial dominated failure, which underscores the crucial role of fiber–matrix interface debonding in the transverse tensile failure of carbon–carbon composites. These uncertainty quantification results facilitate more efficient model calibration and provide a framework for microstructure-informed failure predictions in the face of manufacturing-induced uncertainty.

36 MATERIALS SCIENCE

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE