Search NASASearch

SEARCH · Search NASA

Results for “conformal prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

249 records · Page 3

Predicting Atomistic Transitions with Transformers

Accurate knowledge of the atomistic transition pathways in materials and material surfaces is crucial for many material science problems. However, conventional simulation techniques used to find these transitions are extremely computationally intensive. Even with large-scale, accelerated material simulations, the computational cost constrains the applicable domain in practice. Machine learning models, with the potential to learn the complex emergent behaviors governing atomistic transitions as a fast surrogate model, have great promise to predict transitions with a vastly reduced computational cost. Here, we demonstrate how transformers can be trained to predict atomistic transitions in nano-clusters. We show how we evaluate physical validity of the predictions and how a multitude of additional, different microstates can be generated by slightly varying the data provided to the model.

36 MATERIALS SCIENCE

The Role of Nuclear Data Sensitivities in Prompt α-Eigenvalue Predictions of Delayed Critical Benchmarks

Alpha (α) eigenvalues, which describe the logarithmic time derivative of the neutron population in a multiplying system, are integral to time-dependent behavior and diagnostic applications. However, uncertainties in the evaluated nuclear data can significantly impact the accuracy of transport simulations for such quantities. This work explores the use of machine learning models to predict two key outputs, α-eigenvalues and keff bias, using input features derived from α-eigenvalue sensitivities to nuclear data. The criticality safety benchmark models used in this study come from the International Handbook of Evaluated Criticality Safety Benchmark Experiments. Three models, random forest, XGBoost, and NGBoost, are trained on both energy-resolved and energy-summed α sensitivities. For the α-eigenvalue bias prediction, NGBoost achieved the highest R 2 (0.9476) using energy-resolved features, while XGBoost performed best using summed sensitivities. In contrast, when predicting the keff bias, all the models showed moderate predictive capability (best R 2 ≈ 0.72), as the mapping from the static α-sensitivities to the static keff bias was less direct. SHAP (SHapley Additive exPlanations) analysis was used to interpret the model predictions. Across both prediction tasks, the features associated with neutron capture [H-1 (n, γ)], uranium scattering reactions (such as 235 U elastic/inelastic), and actinide capture/fission reactions (such as 239 Pu and 234 U) were consistently identified as the most impactful. This highlights the key role of specific nuclear reactions and energy ranges in shaping both time-dependent and steady-state criticality behavior. These results demonstrated that α-sensitivities, despite being computed for time-dependent metrics, can provide valuable insights for predicting both α-eigenvalues and the keff bias. Moreover, machine learning models offer a promising pathway for uncovering important nuclear data dependencies and guiding future data evaluation efforts.

Nuclear data

Connectivity, Pathology, and ApoE4 Interactions Predict Longitudinal Tau Spatial Progression and Memory

ABSTRACT Tau pathology spread into neocortex indicates a transition from healthy aging to Alzheimer's disease (AD). Connectivity between tau epicenters and later accumulating regions of cortex has been proposed as a mechanism of tau spread, but how this relationship changes with greater AD pathology burden or genotype is not understood. We investigated tau accumulation in two key regions, precuneus and inferior temporal cortex, using resting state functional connectivity (rsFC) and longitudinal PET imaging from a multicohort sample of cognitively unimpaired older adults. We examined how baseline tau PET, Aβ PET, and ApoE4 genotype status interact with rsFC between hippocampus and these downstream regions to predict rate of tau accumulation in neocortex. We found that the 3‐way interaction between connectivity, baseline tau, and baseline Aβ or ApoE4 status was associated with neocortical tau accumulation in precuneus and inferior temporal cortex. In addition, baseline tau, Aβ, and ApoE4 status also moderated the association between connectivity and rate of memory decline. Together, these results suggest that the extent and distribution of future tau accumulation may be predicted by the interaction of baseline connectivity, AD pathology, and genetic risk.

Neurosciences & Neurology

Wide-ranging predictions of new stable compounds powered by recommendation engines

The computational search for new stable inorganic compounds is faster than ever, thanks to high-throughput density functional theory (DFT). However, stable compound searches remain highly expensive because of the enormous search space and the cost of DFT calculations. To aid these searches, recommendation engines have been developed. We conduct a systematic comparison of the performance of previously developed recommendation engines, specifically ones based on elemental substitution, data mining, and neural network prediction of formation enthalpy. After identifying ways to improve the recommendation engines, we find the neural network to be superior at recommending stable Heusler compounds. Armed with improved recommendation engines, we identify tens of thousands of compounds that are stable at zero temperature and pressure, now available in the Open Quantum Materials Database. We summarize this diverse pool of compounds, including the elusive mixed anion compounds, and two of their many applications: thermoelectricity and solar thermochemical fuel production.

Science & Technology - Other Topics

A Cross‐Linked Flexible Metaferroelectrolyte Regulated by 2D/2D Perovskite Heterostructures for High‐Performance Compact Solid‐State Sodium Batteries

Abstract To address the issues of limited ionic conductivity and poor interface stability at room and low temperatures in solid‐state electrolytes, a robust intrinsic ferroelectrolyte or nanoferroelectrolyte strategy for engineering solid‐state flexible ferroelectric composite electrolytes utilizing strongly coupled intrinsic ion conducting 2D/2D sodium‐rich anti‐perovskite (NaRAP)/ferroelectric perovskite heterostructures is introduced. Herein, highly scalable PVDF‐based metaferroelectrolytes with Na 2.99 Ba 0.005 OCl/Ca 2 Na 2 Nb 5 O 16 − (CNNO − ) nanosheets into a ferroelectric poly(vinylidene fluoride‐co‐hexafluoropropylene) (PVDF‐HFP) matrix, through an in situ cross‐linking and spontaneous bridging method, for compact solid‐state sodium batteries (SSBs), are reported. Benefiting from unique well‐dispersed 3D ferroelectric coupled network and the Na 2.99 Ba 0.005 OCl/CNNO − ‐induced PVDF‐HFP ferroelectric β phase, the Na + flux is regulated, thereby inhibiting Na dendrite growth at the interface. Notably, the optimized PH‐5% NC metaferroelectrolyte exhibits rapid ion transport (1.11 × 10 −4 S cm −1 at 25 °C), a wide electrochemical window (> 4.8V), superior conformal mechanical compatibility, improved flexibility, good elasticity and flame retardancy. The solid‐state Na 3 V 2 (PO 4 ) 3 /PH‐5% NC/Na batteries present a stable cycling performance (remaining 56.4 mAh g −1 after 500 cycles at 1 C) even at 0 °C, potential for cost‐effective, safe, stable and compact SSB energy storage over 600 Wh L −1 , vastly surpassing 365 Wh L −1 of the current commercial sodium‐ion liquid‐electrolyte batteries.

Chemistry

Electronic structure prediction of medium and high entropy alloys across composition space

We propose machine learning (ML) models to predict the electron density — the fundamental unknown of a material’s ground state — across the composition space of concentrated alloys. From this, other physical properties can be inferred, enabling accelerated exploration. A significant challenge is that the number of descriptors and sampled compositions required for accurate prediction grows rapidly with species. To address this, we employ Bayesian Active Learning (AL), which minimizes training data requirements by leveraging uncertainty quantification capabilities of Bayesian Neural Networks. Compared to the strategic tessellation of the composition space, Bayesian-AL reduces the number of training data points by a factor of 2.5 for ternary (SiGeSn) and 1.7 for quaternary (CrFeCoNi) systems. We also introduce easy-to-optimize, body-attached-frame descriptors, which respect physical symmetries while keeping descriptor-vector size nearly constant as alloy complexity increases. Our ML models demonstrate high accuracy and generalizability in predicting both electron density and energy across composition space.

materials science

Self-aligned heterogeneous quantum photonic integration

Integrated quantum photonics holds significant promise for scalable photonic quantum information processing, quantum repeaters, and quantum networks, but its development is hindered by the mismatch between materials hosting high-quality quantum emitters and those compatible with mature photonic technologies. Heterogeneous integration offers a potential solution to this challenge, yet practical implementations have been limited by inevitable insertion losses at material interfaces. Here, we present a self-aligned heterogeneous quantum photonic integration approach that enables near-unity coupling efficiency at the interface. To showcase our approach, we demonstrate Purcell enhancement of a silicon vacancy (SiV) center in diamond induced by a heterogeneous photonic crystal cavity defined by titanium dioxide (TiO 2 ), as well as optical spin control and readout via a TiO 2 photonic circuit. We further show that, when combined with inverse photonic design, our approach enables efficient and broadband collection of single photons from a color center into a heterogeneous waveguide. Our approach is not restricted to SiV centers or TiO 2 ; it has the potential to be broadly applied to integrate diverse solid-state quantum emitters with thin-film photonic devices where conformal deposition is possible. Together, these results establish a practical route to scalable quantum photonic integrated circuits that combine high-quality quantum emitters with technologically mature photonic platforms.

36 MATERIALS SCIENCE

Next-Generation Energy Technologies for Connected and Automated On-Road Vehicles (NEXTCAR) - Predictive Data-Driven Vehicle Dynamics and Powertrain Control: from ECU to the Cloud (Final Scientific/Technical Report)

This project developed and demonstrated a predictive, data-driven vehicle control system designed to improve energy efficiency and driving performance. The team created intelligent self-driving car technology that optimizes fuel and electricity use by proactively planning vehicle actions. By combining Level 4 autonomous driving capabilities with vehicle-to-everything (V2X) connectivity, the system enables vehicles to adjust speed and change lanes in response to traffic signals, surrounding vehicles, and road conditions, reducing unnecessary stops and delays. In testing, the system improved vehicle fuel economy by more than 30% and reduced travel time by approximately 10%, compared to a conventional adaptive cruise control baseline. These results demonstrate the technical effectiveness of using predictive, V2X-enabled strategies, such as traffic light timing and surrounding traffic awareness, to inform real-time vehicle powertrain control and driving behavior. Additionally, a supporting cloud platform was developed to provide dispatch and route recommendations as well as to log vehicle data, demonstrating the economic feasibility of this approach at the fleet level. By optimizing dispatching and routing operations, this technology enables electric fleet operators to use their vehicles more efficiently and reduce reliance on diesel backups, lowering both operating costs and energy consumption. Overall, this project’s technology advances the future of clean, energy-efficient transportation, enabling vehicles and fleets to reduce energy waste, cut costs, and lower emissions through intelligent automation and connectivity.

33 ADVANCED PROPULSION SYSTEMS

A physically interpretable precursor framework for sub-seasonal prediction of Northern Hemisphere flash flourishing

Flash flourishing describes rapid vegetation increases that can quickly reshape land–atmosphere exchanges and impacts on ecosystem, yet its large-scale precursors, circulation context, and sub-seasonal predictability remain poorly understood. Here, we identified onset-stage circulation regimes across northern extratropical latitudes (NEL; >30°N) using 200 and 1000 hPa geopotential height, and examined their regional expressions over eastern Asia, western North America, and Europe. Flash flourishing onset in East Asian was associated with a baroclinic circulation regime and was preceded by a North Atlantic sea surface temperature (SST) precursor at a four-pentad lead. In contrast, onset in western North American and European preferentially occurred under barotropic regimes, preconditioned by Great Plains soil moisture at three-pentad lead and North Atlantic SST at a four-pentad lead, respectively. Ridge regression forecasts revealed regime-dependent sub-seasonal predictability, with mean out-of-sample R 2 exceeding 0.3 up to lead times of two pentads in East Asia, three pentads in western North America, and four pentads in Europe. Together, these findings established a mechanistic and regionally specific framework for anticipating rapid vegetation greening at sub-seasonal timescales.

Kong, Xiangxu [Nanjing Univ. of Information Scienc

Prime Time for Model-Predictive Control? Assessing the Technical and Market Readiness of Advanced Controls in Buildings

Despite three decades of extensive research and field testing that have consistently validated the benefits of Model Predictive Control (MPC) in building applications, the technology has seen limited market adoption. This paper evaluates the readiness of MPC for widespread deployment, showcases recent demonstrations and field tests across diverse building types, including residential, small commercial, large commercial, and campus settings. Our results demonstrate that MPC can optimize system operations to achieve load shifting, minimize curtailment of on-site generation, and reduce energy costs by up to 80 %, while maintaining or improving occupant comfort. We also show that MPC can effectively control large assets, such as MW-sized thermal storage systems, and respond to dynamic pricing signals. However, achieving scale remains difficult due to labor-intensive workflows, reliance on a “PhD-in-the-loop” for MPC design and maintenance, susceptibility to fragile data infrastructure, and persistent workforce education and acceptance barriers. To bridge this gap, we outline a transition from bespoke, labor intensive prototypes toward streamlined, segment-targeted deployment strategies that leverage model templates, semantic tools, and generative AI. By automating control configuration and reducing engineering effort, these recommendations provide a pathway for transforming successful research demonstrations into scalable, market ready solutions for MPC-based controls.

Pritoni, Marco

Microstructure prediction for Ti-22Al-25Nb in laser powder bed fusion

This work presents a physics-informed framework for predicting solidification morphology and defect susceptibility in additively manufactured Ti–22Al–25Nb across a broad processing space. The framework integrates solidification microstructure selection (SMS) analysis with a single-track defect-based printability map to establish a unified methodology linking processing parameters to both interfacial morphology and manufacturability. Thermal gradients G and solidification rates R are first computed using the Thermo-Calc Additive Manufacturing (TC-AM) module, a finite-interface-dissipation (FID) phase-field (PF) model coupled with CALPHAD method is then employed to systematically distinguish planar and dendritic regimes as functions of $G$ and $R$. By superimposing the printability map onto the morphology projections, a comprehensive process–structure framework is obtained. Across most processing conditions, the predicted microstructure is predominantly dendritic, while planar growth emerges only under selected laser power $P$ and scan speed $v$ combinations. In addition to morphology classification, the framework quantifies the dendritic area fraction and introduces a width-based morphology descriptor to characterize the spatial extent of planar/dendritic regions within the melt pool. It provides mechanistic insight into the interplay between solidification physics and defect formation, offering practical guidance for parameter selection and microstructural control in Ti–22Al–25Nb additive manufacturing (AM).

36 MATERIALS SCIENCE

Nuclear–Electronic Orbital General Rate Theory: Predicting Hydrogen Kinetic Isotope Effects in the Deep Tunneling Regime

Hydrogen transfer is a critical component of many chemical and biological processes. The ratio of rate constants for hydrogen and deuterium transfer defines the H/D kinetic isotope effect (KIE), which is a powerful tool for elucidating hydrogen transfer mechanisms. Interpretation of experimental H/D KIEs relies on accurate and affordable computational methods. However, due to their light mass, hydrogen and deuterium can undergo tunneling, which is challenging to describe in multidimensional molecular systems. Herein, we introduce the nuclear–electronic orbital general rate theory (NEO-GRT), which enables the efficient prediction of H/D KIEs based on full-dimensional molecular quantum chemistry calculations. The NEO-GRT approach describes the hydrogen transfer rate constant with a general expression that spans the vibrationally adiabatic and nonadiabatic hydrogen tunneling regimes. The input quantities are computed using NEO density functional theory, which treats the transferring hydrogen or deuterium nucleus quantum mechanically on the same level as the electrons. We investigate two intramolecular proton transfer reactions in organic molecules at temperatures down to 50 K to evaluate the performance of NEO-GRT by comparison to transition state theory and ring-polymer instanton theory. The KIEs computed with NEO-GRT agree with those calculated using ring-polymer instanton theory for the full-dimensional molecular systems at the same level of electronic structure theory. This agreement indicates that NEO-GRT captures the deep hydrogen tunneling effects, in contrast to transition state theory, which neglects such effects. Given its relatively low computational cost, NEO-GRT is a promising approach for predicting H/D KIEs in large organic and organometallic systems.

Hydrogen

Multitask graph neural networks for elastoplastic response prediction in dual-phase polycrystals

Microstructure-sensitive prediction of elastoplastic response remains a recurring bottleneck in multiscale damage and fatigue modeling, where large ensembles of statistically distinct polycrystals are required to quantify variability and extreme-value behavior. In this work, we develop a multitask graph neural network (GNN) surrogate that maps dual-phase ferrite–martensite polycrystal microstructures to Statistical Volume Element (SVE)-level elastoplastic Quantities of Interest (QoIs). Each SVE is represented as a grain-adjacency graph, with node features encoding phase, geometry, and crystallographic orientation, and edge features encoding relative misorientation. A message-passing graph convolution generates node embeddings, which are pooled into a graph representation and passed to a multitask regression head that jointly predicts 10 scalar QoIs and vector-valued stress–strain responses in orthogonal loading directions across multiple martensite volume fractions and SVE sizes. Results show high accuracy for scalar QoIs and strong agreement for full stress–strain trajectories, with population envelopes reproducing both median behavior and finite-SVE variability across compositions and partition scales. A unified model trained on pooled volume-fraction data preserves most within-regime accuracy relative to regime-specific models while also capturing the broader cross-regime variation reflected in the pooled test set. Distributional comparisons further demonstrate that the surrogate preserves heterogeneity under SVE partitioning, enabling statistically consistent block-wise random-field construction for mesoscale analyses. Overall, the proposed grain-graph surrogate provides a practical pathway to accelerate ensemble-based studies of SVE-level constitutive variability in dual-phase polycrystals.

Crystal plasticity

Impact of Crystalline Phases on Low-Activity Waste Glass Durability: Insights from PCT and VHT

During vitrification of nuclear wastes, slow cooling along the container centerline promotes crystalline phase formation, which can alter residual glass composition and reduce chemical durability. This study investigates the effects of crystalline phases on the chemical durability of low-activity waste (LAW) borosilicate glasses using the product consistency test (PCT) and vapor hydration test (VHT) on container centerline cooled (CCC) samples. A preliminary model (R2 = 0.88) was developed to predict CCC PCT responses based on glass composition, PCT data from quenched glasses, and measured crystal fractions. Using the latest LAW glass dataset, the feasibility of predictive modeling is evaluated, limitations in current data and methods are identified, and challenges for improving model accuracy are discussed to guide future data collection and model development.

borosilicate glass

Circumventing data imbalance in magnetic ground state data for magnetic moment predictions

Abstract Magnetic materials play a crucial role in the transition to more sustainable forms of energy and electric vehicles. There is an anticipated shortage in magnetic materials in the future, and as a result there is an urgent need to discover and design new magnetic materials. Computational magnetic material design using density functional theory is daunting because of the challenge in identifying magnetic ground states from a combinatorially large set of possibilities. Machine learning offers a path forward by enabling efficient surrogate models that can more readily enumerate these states, but there is a dearth of training data available, and what is available tends to be imbalanced with too much non-magnetic data. In this work we show that the discrete and previously tackled data imbalance that exists at the level of the magnetic ordering leads to an imbalanced continuous distribution with many zeros when the data is unraveled at the atomic magnetic moment level, which subsequently leads to models with low accuracy for magnetic properties. We mitigate this by using a two-part model framework. Our scheme is able to classify atoms into magnetic and non-magnetic with an F1 score and Matthew’s correlation coefficient (MCC) of ~91% and then to provide an implicit embedding representation that maps directly onto the magnitude of the magnetic moment with a mean absolute error of 0.1 μ B . Beyond screening for new magnetic materials, we demonstrate an additional practical use case of our scheme: the provision of good initial guesses for magnetic moments in first-principles electronic relaxations. Such initialization is shown to lead to faster convergence to configurations that lie closer to the ground state.

Computer Science

Audible Noise Modeling of Hydrogen Release Sonic Hazards in Rail Maintenance Facilities

This study implemented validated literature models to predict audible noise due to pressurized gaseous hydrogen releases through a thermally-activated pressure relief device (TPRD) and attached vent stack. A literature survey discovered limited hydrogen-specific noise prediction models validated by experiments. However, empirical noise prediction models for air flowing through pipes and valves were identified. These empirical models were used to predict noise levels and compared against hydrogen noise data reported in two studies: one experimental study of noise from hydrogen leaking through a pipe and another which modeled hydrogen flowing through a solenoid valve during a fuel cell vehicle refueling. The valve flow model was then applied to predict noise for hydrogen releases through a TPRD. Results show that hydrogen releases through a TPRD can produce harmful noise levels varying from 134 to 150 dB. However, further model validation and additional experimental data are needed to improve prediction confidence and accuracy.

08 HYDROGEN

High-resolution modeling of indoor radon exposure with uncertainty quantification in Utah

Indoor radon accounts for 37% of population-level exposure to ionizing radiation in the United States. However, radon metrics are typically reported at coarse spatial scales, potentially obscuring meaningful local variation. We developed a high-resolution modeling framework to estimate indoor radon concentrations across Utah while explicitly quantifying predictive uncertainty. A total of 19,497 residential radon measurements collected between 2006 and 2017 were combined with environmental and housing characteristics and analyzed using a geospatial neural network that accommodates spatial dependence and nonlinear associations. Predictions were generated on a uniform hexagonal grid at 0.73 km2 resolution (H3 level 8). Out-of-sample predictions aggregated to the H3 level 8 grid showed good agreement with observed concentrations (Pearson r=0.64), while household-level predictions exhibited more moderate agreement (r=0.45). The model produced well-calibrated uncertainty estimates, with 24.1% of held-out observations exceeding the predicted 75th-percentile threshold. Maps of predicted radon concentrations and the probability of exceeding the U.S. EPA action level of 148 Bq/m3 (4 pCi/L) revealed substantial fine-scale spatial heterogeneity that was not apparent in conventional coarse-resolution summaries, with greater local variability observed in densely monitored urban counties than in sparsely sampled regions. High-resolution radon models that explicitly quantify uncertainty provide a useful framework for characterizing the spatial distribution of indoor radon and identifying areas of elevated exceedance risk. These findings highlight the value of fine-scale monitoring data and uncertainty-aware modeling approaches for radon exposure assessment, environmental risk characterization, and radon-related health research.

Wu, Yunhan [ORNL] (ORCID:0000000178842994)

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE