Search NASA⌕ Search

SEARCH · Search NASA

Results for “software structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Software Quality Assurance Plan ANSYS LSDYNA Version 2023R1

ANSYS Inc. develops and markets engineering simulation software and services used in the aerospace, automotive, manufacturing, electronics, biomedical, energy, defense, and many other industries. ANSYS is dedicated to engineering simulation and is the world’s leading software provider. ANSYS was founded in 1970 and is headquartered in Canonsburg, Pennsylvania. ANSYS provides an engineering analysis tool combining structural, thermal, computational fluid dynamics, acoustic and electromagnetic simulation capabilities. ANSYS LS-DYNA is the most used explicit simulation program capable of simulating the response of materials to short periods of severe loading. Its many elements, contact formulations, material models, and other controls can be used to simulate complex models with control over all the details of the problem. ANSYS LS-DYNA has a vast array of capabilities to simulate extreme deformation problems using its explicit solver. Engineers can tackle simulations involving material failure and look at how the failure progresses through a part or through a system. Models with large amounts of parts or surfaces interacting with each other are also easily handled, and the interactions and load passing between complex behaviors are modeled accurately. Using computers with higher numbers of CPU cores can drastically reduce solution times. In addition, many consulting firms and hundreds of universities use ANSYS for analysis, research, and educational purposes. ANSYS is recognized worldwide as one of the most widely used and capable programs of its type. ANSYS has successfully passed over 100 customer quality system audits against American Society of Mechanical Engineers (ASME) NQA-1 and 10 CFR Part 50, Appendix B, since the company was founded, over 60 of which have been since 1997. ANSYS has successfully passed over 100 International Organization for Standardization (ISO) 9001 assessments. ANSYS design analysis software is the first created within a quality system with ISO 9001 certification, which is the internationally accepted quality standard. Product development, testing, maintenance, and support processes also meet the US Nuclear Regulatory Commission’s (NRC’s) quality requirements, as they have for nearly four decades. ANSYS staff perform more than 60,000 software verification tests before releasing each new product. ASME NQA-1-2012 (Subpart 2.7 is specific to software) is the industry- and NRC-accepted approach (consensus standard) for meeting 10 CFR Part 50, Appendix B, requirements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE↗

Python wrapper library and analysis functions for Geotab Altitude API [SWR-24-77]

This software library serves as a Python wrapper for Geotab's Altitude API. It streamlines querying of the API, converts loosely structured API outputs into a standardized tabular data format, and enables analysis of the resulting data tables. It also includes example notebooks showing how to use the library.

Bruchon, Matthew↗

Recent Developments in DFTB+, a Software Package for Efficient Atomistic Quantum Mechanical Simulations

DFTB+ is a flexible, open-source software package developed by its community, designed for fast and efficient atomistic quantum mechanical simulations. It employs various methods that approximate density functional theory (DFT), such as density functional-based tight binding (DFTB) and the extended tight binding (xTB) approach allowing simulations of large systems over extended time scales with reasonable accuracy, while being significantly faster than traditional ab initio methods. In recent years, several new extensions of the DFTB method have been developed and implemented in the DFTB+ program package in order to improve the accuracy and generality of the available simulation results. In this paper, we review those enhancements, show several use case examples and discuss the strengths and limitations of its features.

36 MATERIALS SCIENCE↗

A continuous symmetry breaking measure for finite clusters using Jensen-Shannon divergence

A quantitative measure of symmetry breaking is introduced that allows the quantification of which symmetries are most strongly broken due to the introduction of some kind of defect in a perfect structure. The method uses a statistical approach based on the Jensen-Shannon divergence. The measure is calculated by comparing the transformed atomic density function with its original. Software code is presented that carries the calculations out numerically using Monte Carlo methods. The behavior of this symmetry breaking measure is tested for various cases including finite size crystallites (where the surfaces break the crystallographic symmetry), atomic displacements from high symmetry positions, and collective motions of atoms due to rotations of rigid octahedra. Finally, the approach provides a powerful tool for assessing local symmetry breaking and offers new insights that can help researchers understand how different structural distortions affect different symmetry operations.

atomic & molecular structure↗

Convergent Protocols for Computing Protein–Ligand Interaction Energies Using Fragment-Based Quantum Chemistry

Fragment-based quantum chemistry methods offer a way to sidestep the steep nonlinear scaling of electronic structure calculations so that large molecular systems can be investigated using high-level methods. Here, we use fragmentation to compute protein–ligand interaction energies in systems with several thousand atoms, using a new software platform for managing fragment-based calculations that implements a screened many-body expansion. Convergence tests using a minimal-basis semiempirical method (HF-3c) indicate that two-body calculations, with single-residue fragments and simple hydrogen caps, are sufficient to reproduce interaction energies obtained using conventional supramolecular electronic structure calculations, to within 1 kcal/mol at about 1% of the computational cost. We also demonstrate that the HF-3c results are illustrative of trends obtained with density functional theory in basis sets up to augmented quadruple-ζ quality. Strategic deployment of fragmentation facilitates the use of converged biomolecular model systems alongside high-quality electronic structure methods and basis sets, bringing ab initio quantum chemistry to systems of hitherto unimaginable size. This will be useful for generation of high-quality training data for machine learning applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fidelity Comparison of Time-Bin and Fock State Encoding in Hybrid Quantum Systems Under Channel and Transduction Effects

Future quantum networks are expected to integrate a heterogeneous combination of quantum systems, owing to the diverse advantages offered by different physical platforms in terms of scalability, coherence times, and interfacing capabilities. Within the context of this emerging quantum communication paradigm, this paper presents an analytical comparison of two photonic encoding schemes—time-bin and Fock state (single rail)—in hybrid quantum systems where flying qubits are entangled with stationary matter qubits. We evaluate their resilience against fiber channel and quantum transduction effects by calculating the fidelity of the final states relative to their ideal forms. Employing the characteristic function approach, we derive analytical fidelity expressions and investigate their dependence on parameters such as transmissivity, noise levels, fiber length, and source generation success probability. Additionally, we simulate the scenario with a dedicated QuTiP software implementation to verify the validity of the theoretical models. Our findings reveal that due to its inherent single-mode structure, the Fock state encoding consistently outperforms time-bin encoding in fidelity, as this structure significantly minimizes susceptibility to losses compared to the two-mode nature of the time-bin scheme. This analysis offers valuable insights for future hybrid quantum communication and information processing applications.

Fiorini, Francesco [Pisa U.]↗

SFold v0.1

This is a scientific software package to integrate Small Angle X-ray Scattering (SAXS) experimental data into OpenFold deep learning models to improve protein structure prediction.

Prince, Stephanie [Lawrence Berkeley National Labo↗

Learning Operators for Structure-Informed Surrogate Models

This report summarizes the work performed under the author's two-year John von Neumann LDRD project, which involves the non-intrusive surrogate modeling of dynamical systems with remarkable structural properties. After a brief introduction to the topic, technical accomplishments and project metrics are reviewed including peer-reviewed publications, software releases, external presentations and colloquia, as well as organized conference sessions and minisymposia. The report concludes with a summary of ongoing projects and collaborations which utilize the results of this work.

97 MATHEMATICS AND COMPUTING↗

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La↗

ForceFinder

SAND2025-11750O ForceFinder extends the Structural Dynamics Python Libraries (SDynPy) with comprehensive tools for inverse source estimation (ISE) tasks via frequency response function (FRF) matrix inversion. The software is designed for transfer path analysis and multiple-input/multiple-output (MIMO) vibration control problems. It allows users to estimate sources through various algorithms, from the basic Moore-Penrose pseudo-inverse to statistical learning methods such as Tikhonov regularization via an L-curve and elastic net regularization via an information criterion. ForceFinder uses an object-oriented framework, where all components of the ISE problem—such as FRFs, responses, and transformations—are stored in a "SourcePathReceiver" object. This software can be applied to any noise and vibration problem. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Carter, Steven [Sandia National Lab. (SNL-CA), Liv↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

Terrestrial laser scanning data (Levels 0 and 1) for Pasoh, Malaysia, Sep 2024

This data package contains data from terrestrial laser scanning (TLS) at the Pasoh Forest Reserve, Malaysia. The Pasoh Forest Reserve is a facility of the Forest Research Institute Malaysia, and contains evergreen lowland dipterocarp forest. The Next-Generation Ecosystem Experiments Tropics (NGEE-Tropics) study areas at Pasoh were established to study how different species respond to climatic variation and soil water availability. Two study areas were chosen representing different topography and species. The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree-level characterization of woody structure and leaf area for 12 focal trees with FloraPulse and sap flux sensors, facilitating estimation of woody biomass and leaf area to allow upscaling of water content and transpiration data to the tree-level. Scan positions were not selected to provide consistent data for non-focal trees with the study areas. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

A generalized wind turbine cross section as a reduced-order model to gain insights in blade aeroelastic challenges

In this work, we present an approach to study the aeroelastic stability of a wind turbine by focusing on the dynamics of a blade cross section. We present a methodology to obtain a reduced-order model of the blade dynamics in the form of generalized cross-sectional quantities that approximates the aerodynamic and structural properties of the full blade. The motivation for the work is to gain a physical understanding of the influence of aerodynamic models such as dynamic wake and dynamic stall on the frequency and damping of the structure using a reduced-order model with low computational cost. The model may be coupled to two-dimensional computational fluid dynamics softwares or engineering unsteady airfoil aerodynamics models accounting for dynamic wake and dynamic stall. In the latter case, we can obtain monolithic state-space forms of the aeroelastic system of equations, which simplifies the determination of the modal parameters and therefore the study of stability. The work investigates wind turbines in operation or at standstill, where vortex-induced vibrations and stall-induced vibrations, respectively, might be an issue. The implementation is made available as part of the open-source Python package WELIB and as part of the open-source unsteady aerodynamic driver of OpenFAST.

17 WIND ENERGY↗

Real Vector Framework

SAND2025-11463O The Real Vector Framework (RVF) is a modern and flexible C++ vector math library for developing scientific computing software that involves vector computations. RVF allows an opt-in approach to functionality that parallels the familiar base-class and override structures of object-oriented programming. Users can reuse and customize the code without inheritance entanglements and dynamic dispatch, while enabling seamless interoperability between diverse container types. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

von Winckel, Gregory [Sandia National Lab. (SNL-CA↗

Terrestrial laser scanning data (Levels 0 and 1) from Urban Biogeochemistry Pilot Project sites, Knoxville, Tennessee, Jul 2024 - Jul 2025

This data package contains data from terrestrial laser scanning (TLS) at five urban park sites in Knoxville, Tennessee, USA. All parks include open-grown and/or closed-canopy trees and mixed nearby land use. These study sites were established as part of the Urban Biogeochemistry Pilot Project, which has an overall goal of better understanding how hydrobiogeochemical cycling is altered within the human environment. These five sites represent a gradient of urbanization, and were instrumented to understand hydrological and biogeochemical cycling (e.g., soil moisture, soil physical properties and biogeochemistry, tree transpiration, species type). The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree- and stand-level characterization of woody structure and leaf area. TLS scans were placed to capture the area around trees with sap flow sensors, and as much of a 50 m radius area around the meteorological station as possible given site property limits. Derived products will allow upscaling of water content and transpiration data. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗