Search NASA⌕ Search

SEARCH · Search NASA

Results for “Completeness”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Construction of the damped Ly⁢𝛼 absorber catalog for DESI DR2 Ly⁢𝛼 BAO

We present the Damped Ly⁢𝛼 Toolkit for automated detection and characterization of damped Ly⁢𝛼 absorbers (DLAs) in quasar spectra. Our method uses quasar spectral templates with and without absorption from intervening DLAs to reconstruct observed quasar forest regions. The best-fitting model determines whether a DLA is present while estimating the redshift and HI column density. With an optimized quality cut on detection significance (Δ⁢𝜒$^{2}_{𝑟}$ >0.03), the technique achieves an estimated 80% purity and 79% completeness when evaluated on simulated spectra with S/N>2 that are free of broad absorption lines (BALs). We provide a catalog containing candidate DLAs from the DLA Toolkit detected in DESI DR1 quasar spectra, of which 21 719 were found in S/N>2 spectra with predicted log 10 ⁡(𝑁 𝙷𝙸 )>20.3 and detection significance Δ⁢𝜒$^{2}_{𝑟}$ >0.03. We compare the Damped Ly⁢𝛼 Toolkit to two alternative DLA finders based on a convolutional neural network and Gaussian process models. We present a strategy for combining these three techniques to produce a high-fidelity DLA catalog from DESI DR2 for the Ly⁢𝛼 forest baryon acoustic oscillation measurement. The combined catalog contains 41 152 candidate DLAs with log 10 ⁡(𝑁 𝙷𝙸 )>20.3 from quasar spectra with S/N>2. We estimate this sample to be approximately 85% pure and 79% complete when BAL quasars are excluded.

79 ASTRONOMY AND ASTROPHYSICS↗

How Much Reserve Fuel: Quantifying the Maximal Energy Cost of System Disturbances

Motivated by the design question of additional fuel needed to complete a task in an uncertain environment, this paper introduces metrics to quantify the maximal additional energy used by a control system in the presence of bounded disturbances, compared to a nominal, disturbance-free system. In particular, we consider the task of finite-time stabilization for a linear, time-invariant system. We compare the nominal energy required to achieve this task in the disturbance-free system to the worst-case energy over all feasible disturbances. Solving for the worst-case energy over all disturbances first leads to an optimal control problem with a least-squares solution, and then an infinite-dimensional optimization problem where we derive an upper bound on the solution. The comparison of energies is accomplished using additive and multiplicative metrics, for which we derive bounds. Simulation examples on an ADMIRE fighter jet model demonstrate the practicability of these metrics, and their variation with the distance of the initial condition from the origin and the task completion time.

koopman operator, resilience↗

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

Re-Evaluating Virtual Reality Manipulation Techniques for Precise Alignment of Complex 3D Objects

Prior research has developed a number of manipulation techniques that can achieve precise object placement in virtual reality, but studies of these techniques typically use simple objects. We conducted a study comparing two existing techniques, (AMP-IT and WISDOM), during alignment of objects with complex geometry to evaluate the potential influence of geometric complexity on performance, usability, workload and preference. Our findings indicate that participants had faster completion times and higher trial completion rates with AMP-IT on high-precision alignment tasks, contrary to earlier findings that used simple objects. Yet WISDOM is still preferred and considered more usable, despite increased workload and poorer performance, exposing participants' willingness to trade objective performance for comfort during use.

97 MATHEMATICS AND COMPUTING↗

THERMAL MODELING OF HANFORD CESIUM AND STRONTIUM CANISTERS DURING SIMULATED LOADING

A computational fluid dynamics (CFD) model was built to simulate planned testing of heater assemblies within a canister and overpack for the Hanford Lead Canister (HLC) project. The HLC is a canister storage system that will contain heaters to simulate the decay heat of nuclear material and provide the canister storage system with environmental conditions equivalent to the operating conditions on a dry storage pad. The HLC will be equipped with long-term data collection and monitoring systems to provide an early warning of corrosion, pitting, cracking, or other signs of canister degradation that might threaten the integrity of the containment boundary over the potentially long term of dry storage. An important part of the HLC development is to make pretest numerical predictions for the behavior of the heated canister during the simulated radiolytic decay heat testing, which simulates the dry storage system during loading operations. The simulated radiolytic decay heat test is planned for mid-2024 in a configuration that includes the heater assembly, overpack, and canister, but with the lids removed to allow loading cesium and strontium capsules into the canister. One of the goals of the test is to evaluate the thermal behavior of the canister and overpack assembly in the ambient air of the test facility, which will provide data critical to validating the thermal models and understanding how the HLC will perform as a system once deployed. To best approximate real-world conditions, the CFD model includes the full air volume of the mock-up truck bay the heated canister test will be performed in, enabling detailed investigation of how the heated canister affects airflow around it. Rigorous pre-deployment testing of the complete HLC cask and canister system is intended to be completed before the HLC is deployed in the 2028 timeframe. This study presents the pre-test temperature predictions of the simulated radiolytic decay heat test. A description of the heater assembly, canister, and overpack system is presented. The model was developed with the commercial CFD software STAR-CCM+. An uncertainty analysis was run with the CFD model to determine the uncertainty in the temperature predictions and provide a range over which the predicted temperatures are expected to vary. The uncertainty analysis was preformed by coupling STAR-CCM+ with the software Dakota, which provides advanced parametric analyses, including quantification of margins and uncertainty with computational models. This work is expected to provide insight into SNF canister behavior.

Carpenter-Graffy, Dina E.↗

3D reconstruction and neural rendering for adversarial machine learning

While evasion attacks on computer vision systems have been widely studied, creating attacks that remain effective under significant changes in viewpoint continues to be challenging. Traditional approaches often rely on affine transformations of images, but these approaches degrade at larger perspective shifts and often produce unrealistic or ineffective perturbations. Recent methods use differentiable renderers to improve viewpoint robustness, but they typically depend on manually constructed 3D models. We introduce a semi-automated pipeline that generates physically printable and perspective-invariant adversarial patches using only a small set of 2D images. Our method integrates 3D reconstruction, neural rendering, adversarial patch optimization, and an object detection victim model into a unified workflow. We use 2D Gaussian Splatting for high fidelity mesh reconstruction and FlexPara for surface parameterization that produces texture maps suitable for patch editing. Together, these components form a fully differentiable pipeline in PyTorch3D that links texture modification to model outputs, enabling efficient optimization of patches that remain effective across many viewpoints. The complete process, from image capture to patch printing and physical evaluation, can be completed within a few hours. We demonstrate the effectiveness of the resulting patches through attacks on the YOLOv8 object detection model and discuss remaining challenges and opportunities for improving robustness and scalability.

Singhvi, Vivaan [ORNL] (ORCID:0009000586288221)↗

Leptothrix ochracea genomes reveal potential for mixotrophic growth on Fe(II) and organic carbon

ABSTRACT Leptothrix ochracea creates distinctive iron-mineralized mats that carpet streams and wetlands. Easily recognized by its iron-mineralized sheaths, L. ochracea was one of the first microorganisms described in the 1800s. Yet it has never been isolated and does not have a complete genome sequence available, so key questions about its physiology remain unresolved. It is debated whether iron oxidation can be used for energy or growth and if L. ochracea is an autotroph, heterotroph, or mixotroph. To address these issues, we sampled L. ochracea -rich mats from three of its typical environments (a stream, wetlands, and a drainage channel) and reconstructed nine high-quality genomes of L. ochracea from metagenomes. These genomes contain iron oxidase genes cyc2 and mtoA, showing that L. ochracea has the potential to conserve energy from iron oxidation. Sox genes confer potential to oxidize sulfur for energy. There are genes for both carbon fixation (RuBisCO) and utilization of sugars and organic acids (acetate, lactate, and formate). In silico stoichiometric metabolic models further demonstrated the potential for growth using sugars and organic acids. Metatranscriptomes showed a high expression of genes for iron oxidation; aerobic respiration; and utilization of lactate, acetate, and sugars, as well as RuBisCO, supporting mixotrophic growth in the environment. In summary, our results suggest that L. ochracea has substantial metabolic flexibility. It is adapted to iron-rich, organic carbon-containing wetland niches, where it can thrive as a mixotrophic iron oxidizer by utilizing both iron oxidation and organics for energy generation and both inorganic and organic carbon for cell and sheath production. IMPORTANCE Winogradsky's observations of L. ochracea led him to propose autotrophic iron oxidation as a new microbial metabolism, following his work on autotrophic sulfur-oxidizers. While much culture-based research has ensued, isolation proved elusive, so most work on L. ochracea has been based in the environment and in microcosms. Meanwhile, the autotrophic Gallionella became the model for freshwater microbial iron oxidation, while heterotrophic and mixotrophic iron oxidation is not well-studied. Ecological studies have shown that Leptothrix overtakes Gallionella when dissolved organic carbon content increases, demonstrating distinct niches. This study presents the first near-complete genomes of L. ochracea , which share some features with autotrophic iron oxidizers, while also incorporating heterotrophic metabolisms. These genome, metabolic modeling, and transcriptome results give us a detailed metabolic picture of how the organism may combine lithoautotrophy with organoheterotrophy to promote Fe oxidation and C cycling and drive many biogeochemical processes resulting from microbial growth and iron oxyhydroxide formation in wetlands.

59 BASIC BIOLOGICAL SCIENCES↗

An exopolysaccharide pathway from a freshwater Sphingomonas isolate

Bacteria embellish their cell envelopes with a variety of specialized polysaccharides. Biosynthesis pathways for these glycans are complex, and final products vary greatly in their chemical structures, physical properties, and biological activities. This tremendous diversity comes from the ability to arrange complex pools of monosaccharide building blocks into polymers with many possible linkage configurations. Due to the complex chemistry of bacterial glycans, very few biosynthetic pathways have been defined in detail. As part of an initiative to characterize novel polysaccharide biosynthesis enzymes, we isolated a bacterium from Lake Michigan called Sphingomonas sp. LM7 that is proficient in exopolysaccharide (EPS) production. We identified genes that contribute to EPS biosynthesis in LM7 by screening a transposon mutant library for colonies displaying altered colony morphology. A gene cluster was identified that appears to encode a complete wzy/wzx-dependent polysaccharide assembly pathway. Deleting individual genes in this cluster caused a non-mucoid phenotype and a corresponding loss of EPS secretion, confirming the role of this gene cluster in polysaccharide production. We extracted EPS from LM7 cultures and determined that it contains a linear chain of 3- and 4-linked glucose, galactose, and glucuronic acid residues. Finally, we show that the EPS pathway in Sphingomonas sp. LM7 diverges from that of sphingan-family EPSs and adhesive polysaccharides such as the holdfast that are present in other Alphaproteobacteria. Our approach of characterizing complete biosynthetic pathways holds promise for engineering polysaccharides with valuable properties.

59 BASIC BIOLOGICAL SCIENCES↗

New horizons in the holographic conformal phase transition

Abstract We describe 5D dynamical cosmological solutions of the stabilized holographic dilaton and their role in completion of the conformal phase transition. This analysis corresponds, via the AdS/CFT dictionary, to a study of out-of-equilibrium dynamics where trajectories of the dilaton do not depend solely on thermodynamic quantities in the early universe, but have sensitivity also to initial conditions. Unlike the well-studied thermal transition, which requires quantum tunneling of an infrared brane through the surface of an AdS-Schwarzschild horizon, our approach instead invokes an early epoch in which the cosmology is fully 5-dimensional, with highly relativistic brane motion and with Rindler horizons obscuring the infrared brane at early times. In this context, we demonstrate the existence of a large class of natural initial conditions that seed trajectories where the brane simply passes through the Rindler horizon and into the basin of attraction of the stabilized dilaton potential. This corresponds to successful completion of the phase transition without sacrificing perturbativity of the 5D theory.

Physics↗

Review of Particle Physics

The Review summarizes much of particle physics and cosmology. Using data from previous editions, plus 3,200 new measurements from 903 papers, we list, evaluate, and average measured properties of gauge bosons and the recently discovered Higgs boson, leptons, quarks, mesons, and baryons. We summarize searches for hypothetical particles such as supersymmetric particles, heavy bosons, axions, dark photons, etc. Particle properties and search limits are listed in Summary Tables. We give numerous tables, figures, formulae, and reviews of topics such as Higgs Boson Physics, Supersymmetry, Grand Unified Theories, Neutrino Mixing, Dark Energy, Dark Matter, Cosmology, Particle Detectors, Colliders, Probability and Statistics. Most of the 118 reviews are updated, including many that are heavily revised.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗

Using a Large Language Model as a Building Block to Generate Usable Validation and Verification Suite for OpenMP

In the HPC area, both hardware and software move quickly. Often new hardware is developed and deployed, the corresponding software stack, including compilers and other tools, are under active development while leading edge software developers are working to port and tune their applications, all at the same time. While the software ecosystem is in flux, one of the key challenges for users is obtaining insight into the state of implementation of key features in the programming languages and models their applications are using – whether they have been implemented, and whether the implementation conforms to the specification, especially for newly implemented features (less tested by widespread use). OpenMP is one of the most prominent shared memory programming models used for on-node programming in HPC. With the shift towards accelerators (such as GPUs and FPGAs) and heterogeneous programming OpenMP features are getting more complex. It is natural to ask whether generative AI approaches, and large language models (LLMs) in particular, can help in producing validation and verification test suites to allow users better and faster insights into the availability and correctness of OpenMP features of interest. In this work, we explore the use of ChatGPT-4 to generate a suite of tests for OpenMP features. We have chosen a set of directives and clauses, a total of 78 combinations, which first appeared in OpenMP 3.0 (released in May 2008) but are also relevant for accelerators. We prompted ChatGPT to generate tests in the C and Fortran languages, for both host (CPU) and device (accelerator). On the Summit super-computer using the GNU implementation, we found that, of the 78 generated tests 67 C tests and 43 Fortran tests compiled successfully and fewer than those executed to completion. On further analysis we show that not all generated tests are valid. We document the process, results, and provide detailed analysis regarding the quality of tests generated. With the aim of providing input to a production quality validation and verification suite, we manually implement the corrections required to make the tests valid according to the current OpenMP specification. We quantify this effort as small, medium, or large, and record the lines of code changed to correct the invalid tests. With the corrected tests we validate recent implementations from HPE, AMD, and GNU on the Frontier supercomputer. Our experiment and subsequent analysis show that although LLMs are capable of producing HPC specific codes, they are limited by their understanding of the deeper semantics and restrictions of programming models such as OpenMP. Unsurprisingly more commonly used features have better support, while some OpenMP 3.0 directives such as sections and tasking are not universally supported on accelerators. We demonstrate that successful compilation and execution to completion are inadequate metrics for evaluating generated code and that, at this time, commodity LLMs require expert intervention for code verification. This points to gaps in the training data that is currently available for HPC. We demonstrate that with "small" effort 37% of generated invalid C tests and 63% of generated invalid Fortran tests could be corrected. This improves productivity of test generation as we circumvent writing from scratch and the common programming errors associated with it.

Pophale, Swaroop [ORNL] (ORCID:0000000185446367)↗

Unraveling the Structure and Composition of Li 4 Mn 2 O 4.5 (Li 2 O·Li 0.667 Mn 1.333 O 2 ) Electrodes for Lithium Batteries Using a High-Temperature Synthesis Approach

This paper addresses the debate about the composition and structure of a lithium-rich manganese oxide electrode with a fully disordered rock salt component, Li 4 Mn 2 O 5 (or Li 2 O·2LiMnO 2 ), first reported by Freire et al. in 2016; it is typically prepared by a high-energy ball milling procedure. It has now been demonstrated that, when prepared at 800°C, the formula of this compound is Li 4 Mn 2 O 4.5 , alternatively Li 2 O·Li 0.667 Mn 1.333 O 2 , or close thereto. The cubic, disordered Li 0.667 Mn 1.333 O 2 (or Li 0.333 Mn 0.667 O) rock salt component, in which the manganese ions adopt an average oxidation state of 2.5+, transforms to a clearly-defined spinel configuration during electrochemical cycling. The electrochemical activation process that occurs during the initial charge reaction includes the oxidation of the manganese ions by oxygen released by the Li 2 O component between 4.5 and 4.6 V. In complete contrast, nickel- and nickel-cobalt-substituted electrodes, such as Li 2 O·2LiMn 0.5 Ni 0.5 O 2 (Li 4 MnNiO 5 ) and Li 2 O·2LiMn 0.475 Ni 0.475 Co 0.050 O 2 (Li 4 Mn 0.95 Ni 0.95 Co 0.10 O 5 ), in which the manganese ions adopt a tetravalent state, have completely disordered rock salt components that are electrochemically inactive.

25 ENERGY STORAGE↗

Statistical Design of Experiments Enables Rapid Exploration of Perfluorobutane Sulfonate Degradation

The phase-out of long-chain PFAS has led to the proliferation of highly recalcitrant short-chain analogs, and mineralization technologies for these short-chain PFAS are needed to mitigate their deleterious effects on environmental and human health. In this study, we utilize a statistical design of experiments, specifically, response surface methodology, to rapidly evaluate the electrochemical degradation of the short-chain PFAS, perfluorobutane sulfonate (PFBS). We evaluate the impacts of the three primary electrochemical parameters (concentration of PFBS, concentration of supporting electrolyte, and applied current) over multiple orders of magnitude on the three primary reaction outcomes of electrochemical PFBS degradation (incomplete PFBS decomposition, complete PFBS mineralization as fluorine, and anodic energy consumption). Our results correspond with literature and clearly identify the well-known tradeoff between energy consumption and complete mineralization. Intriguingly, partial PFBS decomposition and energy consumption demonstrate nonlinear dependencies in the current/supporting electrolyte concentration space and the current/PFBS concentration space, respectively. These findings highlight the utility of the response surface methodology model to efficiently interrogate a large parameter space, identifying both common results and less-obvious interactions between electrochemical parameters and their influences on reaction outcomes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Modified Data Collection And Analysis Codes Of Using Tcm (thermal Conductivity Microscope) To Measure Thermal Conductivity And Diffusivity

The "data collection" basically involves setting up the thermal wave frequency, laser scan distance, and other parameters related to the experimental setup. The modification of this code is minor and the details of this code can be found in the earlier patent ("thermal conductivity microscope"). The "data analysis" instead, replaces the simplified analytical model by a more complete analytical model, and used a "thermoquadruple" method to solve the analytical model. The efficiency is orders of magnitude improved and the accuracy is also better. Meanwhile, the previous model can only handle a two-layer sample structure. The new, complete model can handle materials with multiple layers (any given number), which is necessary to handle post ion irradiated materials.

Hua, Zilong [Idaho National Laboratory (INL), Idah↗

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor↗

Early-stage Testing of Thermal Power Dispatch Simulator for Subjective Mental Workload and Evolution Time

The excess thermal energy produced by nuclear power plants (NPPs) during low electricity demands can be utilized in industrial processes, such as hydrogen production, through a thermal power dispatch (TPD) system. Initial testing of the first iteration of a single-train TPD design with a manual control mode at the Idaho National Lab (INL) revealed a high operator workload and degraded control capability. The current study evaluated the impact of an enhanced dual-train TPD design on operators’ subjective mental workload and evolution task-time while completing two operating scenarios in manual and automatic control modes. The results showed no statistically significant difference between participants’ mental workload using both control modes. Evolution time in automatic control mode took a shorter time than in manual control, with participants completing all evolutions in less than the 10-min set as the design specification limit. The shorter evolution time is discussed within the context of plant safety and operational efficiency.

Gideon, Olugbenga↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗