Search NASA⌕ Search

SEARCH · Search NASA

Results for “Task based parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

222 records · Page 13

Configuration and Projected Capabilities of the Common Habitat Medical Care Facility

The Common Habitat is a large, long-duration habitat being explored as part of a conceptual study (not an active NASA program) that uses an SLS core stage Liquid Oxygen (LOX) tank as its primary structure. It is intended for use on the Moon as part of a permanently occupied outpost, on Mars as part of an outpost that will be occupied for hundreds of days at a time, and in deep space as part of the Deep Space Exploration Vehicle where it will support crewed missions up to 1200 days in duration. A study of internal orientation and crew size resulted in a Common Habitat configuration sized for a crew of eight with a three-deck horizontal orientation. Additional work outside the scope of this paper is developing a vertical translation system, a crew mobility aids system based on wearable gecko-derived grippers, and a crew seating/restraint system. These systems are all assumed for use in conjunction with the Medical Care Facility, which is needed to maintain crew well-being during these missions, where distance from Earth precludes the possibility of evacuation to Earth. This paper describes recent improvements in the Common Habitat Medical Care Facility and associated benefits for crew survivability in long duration missions beyond Earth orbit. These improvements were made with the assistance of a NASA Pathways intern whose experience includes a tour of duty in Afghanistan as an Army combat medic with the 691st GHOST-T, attached to the 1st and 7th US Special Forces Groups as part of Operation Freedom’s Sentinel, where he helped provide far-forward surgical capabilities in austere combat environments. The initial baseline Medical Care Facility was developed working in conjunction with University of Houston Space Architecture graduate students. The facility was placed on the upper deck of the Common Habitat in a location that provided privacy, operational volume, and was close to the vertical translation pathway. The notional outfitting repurposed component CAD models from unrelated studies and notionally indicated a level of care roughly equivalent to that aboard the International Space Station. The CAD modeling provided notional stowage volumes, a deployable surface, some fixed equipment, an ultrasound, and a potentially reconfigurable treatment table. While this facility is clearly a competent arrangement, it was desired to leverage available expertise and upgrade the station given the vast distances from Earth to be experienced by the Common Habitat. Key driving requirements applied to the upgrade included to provide Medical Level of Care V, offer enhanced telemedicine capabilities, provide patient physical accommodation, provide caregiver access to the patient from all sides, include sliding pocket doors for access to hygiene and to the Vertical Translation System, and to add any additional capability possible for the best achievable medical care. The first step in the facility upgrade was to quantify the current medical inventory on the International Space Station and ensure that sufficient stowage volume was present for this purpose. To that end, the ISS medical kits were reviewed, and eight full size mid deck lockers were placed in the facility. A number of additional devices were also added, based on the intern’s combat medic experience. Also, two fixed shelves and one horizontal work surface were added to the Medical Care Facility, with the shelves providing storage space for the additional devices and the work surface providing a location for the caregiver to work or stage equipment. Four display monitors were added to the wall above the horizontal work surface, supporting data display, telemedicine, conferencing, or other needs. The existing treatment table was replaced with a mobile surgical stretcher-chair. Two additional doors were added to the Medical Care Facility. One leads directly to the hygiene compartment, allowing it to support medical operations in addition to providing galley/wardroom support. The other door leads directly into the Vertical Translation System. The wall adjacent to the subsystems bay was moved, adding additional volume to the Medical Care Facility. This improved caregiver access to the patient and allowed for a larger number of caregivers to be present. It also provided options for relocation of support equipment relative to the patient as needed. In the upgraded Medical Care Facility, the Surgical Stretcher-Chair and the Vertical Translation System can work together to provide incapacitated crew member transport from a site of injury on any deck of the Common Habitat to the Medical Care Facility. It can also support patient treatment in a variety of positions including a variety of sitting postures and a supine posture at a variety of pitch angles. The facility can also support caregiver office work for review of examination results, private consultation, inventory and maintenance, and a variety of other purposes. A forward activity will be to conduct evaluations of the Medical Care Facility with different medical scenarios. Additionally, ambient and task lighting selections remain as forward work. The eight mid deck lockers can be augmented to use as portable equipment carts, similar to a manner in which maintenance facility stowage was used as portable carts during the NASA Desert Research and Technology Studies in the Constellation Program. Trash accommodation will also need forward work to assess, including provision for wet trash, dry trash, and biological waste. It will be important to assess a redesign of the surgical stretcher-chair. The commercial version used in the upgrade can only enable vertical translation in the seated configuration, requiring the patient to bend both hips and knees. A possible redesign of the chair will allow for vertical translation without requiring any bending at the hip or knees. Also, the commercial version is wheeled, making it mobile in gravity but unanchored in microgravity. Work will be needed to adapt the chair for gravity-independent performance. The hygiene compartment can be redesigned for dual-use medical scrub and galley handwash facility. Pending sufficient volume, it may also be possible to place sanitation equipment in this location to clean medical tools. Finally, most space architectures have never allowed for more than one incapacitated crew member, but several scenarios could potentially injure two or more crew in the same incident. This facility could be assessed to determine its present ability to address two or more injured crew in parallel and determine the potential upper limit for number of treatable crew in a multi-crew injury scenario, or to treat polytrauma of a single patient.

Habitat↗

Configuration and Projected Capabilities of the Common Habitat Medical Care Facility

The Common Habitat is a large, long-duration habitat being explored as part of a conceptual study (not an active NASA program) that uses an SLS core stage Liquid Oxygen (LOX) tank as its primary structure. It is intended for use on the Moon as part of a permanently occupied outpost, on Mars as part of an outpost that will be occupied for hundreds of days at a time, and in deep space as part of the Deep Space Exploration Vehicle where it will support crewed missions up to 1200 days in duration. A study of internal orientation and crew size resulted in a Common Habitat configuration sized for a crew of eight with a three-deck horizontal orientation. Additional work outside the scope of this paper is developing a vertical translation system, a crew mobility aids system based on wearable gecko-derived grippers, and a crew seating/restraint system. These systems are all assumed for use in conjunction with the Medical Care Facility, which is needed to maintain crew well-being during these missions, where distance from Earth precludes the possibility of evacuation to Earth. This paper describes recent improvements in the Common Habitat Medical Care Facility and associated benefits for crew survivability in long duration missions beyond Earth orbit. These improvements were made with the assistance of a NASA Pathways intern whose experience includes a tour of duty in Afghanistan as an Army combat medic with the 691st GHOST-T, attached to the 1st and 7th US Special Forces Groups as part of Operation Freedom’s Sentinel, where he helped provide far-forward surgical capabilities in austere combat environments. The initial baseline Medical Care Facility was developed working in conjunction with University of Houston Space Architecture graduate students. The facility was placed on the upper deck of the Common Habitat in a location that provided privacy, operational volume, and was close to the vertical translation pathway. The notional outfitting repurposed component CAD models from unrelated studies and notionally indicated a level of care roughly equivalent to that aboard the International Space Station. The CAD modeling provided notional stowage volumes, a deployable surface, some fixed equipment, an ultrasound, and a potentially reconfigurable treatment table. While this facility is clearly a competent arrangement, it was desired to leverage available expertise and upgrade the station given the vast distances from Earth to be experienced by the Common Habitat. Key driving requirements applied to the upgrade included to provide NASA’s Medical Level of Care V, offer enhanced telemedicine capabilities, provide patient physical accommodation, provide caregiver access to the patient from all sides, include sliding pocket doors for access to hygiene and to the Vertical Translation System, and to add any additional capability possible for the best achievable medical care. The first step in the facility upgrade was to quantify the current medical inventory on the International Space Station and ensure that sufficient stowage volume was present for this purpose. To that end, the ISS medical kits were reviewed, and eight full size mid deck lockers were placed in the facility. A number of additional devices were also added, based on the co-author’s combat medic experience. Also, two fixed shelves and one horizontal work surface were added to the Medical Care Facility, with the shelves providing storage space for the additional devices and the work surface providing a location for the caregiver to work or stage equipment. Four display monitors were added to the wall above the horizontal work surface, supporting data display, telemedicine, conferencing, or other needs. The existing treatment table was replaced with a mobile surgical stretcher-chair. Two additional doors were added to the Medical Care Facility. One leads directly to the hygiene compartment, allowing it to support medical operations in addition to providing galley/wardroom support. The other door leads directly into the Vertical Translation System. The wall adjacent to the subsystems bay was moved, adding additional volume to the Medical Care Facility. This improved caregiver access to the patient and allowed for a larger number of caregivers to be present. It also provided options for relocation of support equipment relative to the patient as needed. In the upgraded Medical Care Facility, the Surgical Stretcher-Chair and the Vertical Translation System can work together to provide incapacitated crew member transport from a site of injury on any deck of the Common Habitat to the Medical Care Facility. It can also support patient treatment in a variety of positions including a variety of sitting postures and a supine posture at a variety of pitch angles. The facility can also support caregiver office work for review of examination results, private consultation, inventory and maintenance, and a variety of other purposes. A forward activity will be to conduct evaluations of the Medical Care Facility with different medical scenarios. Additionally, ambient and task lighting selections remain as forward work. The eight middeck lockers can be augmented to use as portable equipment carts, similar to a manner in which maintenance facility stowage was used as portable carts during the NASA Desert Research and Technology Studies in the Constellation Program. Trash accommodation will also need forward work to assess, including provision for wet trash, dry trash, and biological waste. It will be important to assess a redesign of the surgical stretcher-chair. The commercial version used in the upgrade can only enable vertical translation in the seated configuration, requiring the patient to bend both hips and knees. A possible redesign of the chair will allow for vertical translation without requiring any bending at the hip or knees. Also, the commercial version is wheeled, making it mobile in gravity but unanchored in microgravity. Work will be needed to adapt the chair for gravity-independent performance. The hygiene compartment can be redesigned to serve both as a medical scrub facility and for galley hand washing. Pending sufficient volume, it may also be possible to place sanitation equipment in this location to clean medical tools. Finally, most space architectures have never allowed for more than one incapacitated crew member, but several scenarios could potentially injure two or more crew in the same incident. This facility could be assessed to determine its present ability to address two or more injured crew in parallel and determine the potential upper limit for number of treatable crew in a multi-crew injury scenario, or to treat polytrauma of a single patient.

Common Habitat↗

Current Status of Martian Moons eXploration (MMX) Contamination Control and Curation Activity

Martian Moons eXploration (MMX) is a sample return mission from the Martian moon Phobos. The MMX spacecraft is scheduled to launch in 2026 and return to Earth in 2031. The main science goals of MMX are “to reveal the origin of the Martian moons and make progress in the understanding of planetary system formation and material transport in the solar system, and to observe processes that impact the circumplanetary and surface environments of Mars”. MMX has two sampling systems: coring (C)-sampler and pneumatic (P)-sampler and plans to bring back >10 g of Phobos sample. The retuned sample in the sample capsule will be transferred to the curation facility in ISAS/JAXA for sample curation and subsequent sample analysis. Contamination control of the sample return mission requires special care to prevent terrestrial contamination to the spacecraft, which would ruin the scientific value of the returned sample. Retaining the pristineness of the retuned sample is an important task of the MMX Curation and Sampler Science teams. The basis of the contamination control is (1) to minimize and understand the nature and amount of contaminants, (2) to perform contamination assessment and evaluate the effect of contaminants in the spacecraft on the retuned sample, (3) to employ a contamination knowledge (CK) material coupon in the spacecraft to identify the contaminants in the returned samples. In the MMX contamination control plan, the allowable contamination level for each contaminant is carefully defined. They are mostly set to be 1/1000 of the expected amount of each material in the returned sample and are divided into two main categories: organic and inorganic. The allowable atmospheric leakage rate to the sample container is also defined. The allowable contamination level of the organic materials is based on the composition of carbonaceous chondrites. The target contaminants are amino acids, aliphatic and aromatic hydrocarbons, carboxylic acids, etc. In case of the inorganic materials, the target contaminants are important elements to permit distinguishing the origin of the Martian moon by nucleosynthetic isotope anomalies (Cr, Ti, and Mo) and to reveal the evolution of the Martian moon by chronology (Hf, W, U, Pb, Rb, Sr, Sm, and Nd). The key instrument of contamination control in the sample return mission is the sampler system. The C-sampler has been developed by JAXA and the P-sampler was provided by Honeybee/NASA. In MMX, materials used in the two samplers (C- and P- sampler) were carefully selected to avoid potential contamination from the design stage of the system. The individual parts of the C-sampler FM (Flight Model) were thoroughly cleaned at the curation facility in ISAS/JAXA by the full-course cleaning procedure, which is an ultrasonic cleaning with organic solvents and ultrapure water in several steps. The equivalent level of cleaning was also carried out on the P-Samper FM as well by Honeybee Robotics in the USA. Now, MMX is in the critical phase for contamination control called ATLO: Assembly, Test, and Launch Operations. During the ATLO phase, sampler FM is constantly purged with nitrogen gas and maintained at positive pressure to prevent environmental contamination. The surrounding environments of the sampler FM are also simultaneously monitored using the CK Monitoring Coupon Set, which consists of several witness materials such as a glass petri dish, sapphire glass disk, and carbon adhesive tape (Figure 1). The detailed environmental assessment of each clean room used for the assembly and test of the sampler FM has also been conducted. This assessment includes microbial analysis, which was performed for OSIRIS-REx. Regarding the sample recovery and sample curation, we have started the designing of Sample Container Disassembling Instrument for the sample recovery from the sample container and the MMX curation chamber for sample curation. The curation protocol for the Phobos returned sample has also been discussed by the MMX Sample Analysis Working Team (SAWT). The MMX curation protocol consists of three phases: (1) quick analysis, (2) pre-basic characterization, and (3) basic characterization. (1) is extraction of the sample gas from the sample container and analysis by mass spectrometry, (2) is observation in bulk level, and (3) is observation in grain level and allocation of the sample aliquots. In parallel with the curation protocol, the returned sample undergoes preliminary examination for scientific investigations to achieve science goals. In addition, the CK witness plates made of sapphire glass are on board the sampler system. The CK witness plates will be recovered from the sampler system and analyzed by SAWT for the assessment of in-flight contamination.

Haruna Sugahara↗

Use Computer-Aided Tools to Parallelize Large CFD Applications

Porting applications to high performance parallel computers is always a challenging task. It is time consuming and costly. With rapid progressing in hardware architectures and increasing complexity of real applications in recent years, the problem becomes even more sever. Today, scalability and high performance are mostly involving handwritten parallel programs using message-passing libraries (e.g. MPI). However, this process is very difficult and often error-prone. The recent reemergence of shared memory parallel (SMP) architectures, such as the cache coherent Non-Uniform Memory Access (ccNUMA) architecture used in the SGI Origin 2000, show good prospects for scaling beyond hundreds of processors. Programming on an SMP is simplified by working in a globally accessible address space. The user can supply compiler directives, such as OpenMP, to parallelize the code. As an industry standard for portable implementation of parallel programs for SMPs, OpenMP is a set of compiler directives and callable runtime library routines that extend Fortran, C and C++ to express shared memory parallelism. It promises an incremental path for parallel conversion of existing software, as well as scalability and performance for a complete rewrite or an entirely new development. Perhaps the main disadvantage of programming with directives is that inserted directives may not necessarily enhance performance. In the worst cases, it can create erroneous results. While vendors have provided tools to perform error-checking and profiling, automation in directive insertion is very limited and often failed on large programs, primarily due to the lack of a thorough enough data dependence analysis. To overcome the deficiency, we have developed a toolkit, CAPO, to automatically insert OpenMP directives in Fortran programs and apply certain degrees of optimization. CAPO is aimed at taking advantage of detailed inter-procedural dependence analysis provided by CAPTools, developed by the University of Greenwich, to reduce potential errors made by users. Earlier tests on NAS Benchmarks and ARC3D have demonstrated good success of this tool. In this study, we have applied CAPO to parallelize three large applications in the area of computational fluid dynamics (CFD): OVERFLOW, TLNS3D and INS3D. These codes are widely used for solving Navier-Stokes equations with complicated boundary conditions and turbulence model in multiple zones. Each one comprises of from 50K to 1,00k lines of FORTRAN77. As an example, CAPO took 77 hours to complete the data dependence analysis of OVERFLOW on a workstation (SGI, 175MHz, R10K processor). A fair amount of effort was spent on correcting false dependencies due to lack of necessary knowledge during the analysis. Even so, CAPO provides an easy way for user to interact with the parallelization process. The OpenMP version was generated within a day after the analysis was completed. Due to sequential algorithms involved, code sections in TLNS3D and INS3D need to be restructured by hand to produce more efficient parallel codes. An included figure shows preliminary test results of the generated OVERFLOW with several test cases in single zone. The MPI data points for the small test case were taken from a handcoded MPI version. As we can see, CAPO's version has achieved 18 fold speed up on 32 nodes of the SGI O2K. For the small test case, it outperformed the MPI version. These results are very encouraging, but further work is needed. For example, although CAPO attempts to place directives on the outer- most parallel loops in an interprocedural framework, it does not insert directives based on the best manual strategy. In particular, it lacks the support of parallelization at the multi-zone level. Future work will emphasize on the development of methodology to work in a multi-zone level and with a hybrid approach. Development of tools to perform more complicated code transformation is also needed.

Jin, H.↗

Investigating the Simulink Auto-Coding Process

Model based program design is the most clear and direct way to develop algorithms and programs for interfacing with hardware. While coding "by hand" results in a more tailored product, the ever-growing size and complexity of modern-day applications can cause the project work load to quickly become unreasonable for one programmer. This has generally been addressed by splitting the product into separate modules to allow multiple developers to work in parallel on the same project, however this introduces new potentials for errors in the process. The fluidity, reliability and robustness of the code relies on the abilities of the programmers to communicate their methods to one another; furthermore, multiple programmers invites multiple potentially differing coding styles into the same product, which can cause a loss of readability or even module incompatibility. Fortunately, Mathworks has implemented an auto-coding feature that allows programmers to design their algorithms through the use of models and diagrams in the graphical programming environment Simulink, allowing the designer to visually determine what the hardware is to do. From here, the auto-coding feature handles converting the project into another programming language. This type of approach allows the designer to clearly see how the software will be directing the hardware without the need to try and interpret large amounts of code. In addition, it speeds up the programming process, minimizing the amount of man-hours spent on a single project, thus reducing the chance of human error as well as project turnover time. One such project that has benefited from the auto-coding procedure is Ramses, a portion of the GNC flight software on-board Orion that has been implemented primarily in Simulink. Currently, however, auto-coding Ramses into C++ requires 5 hours of code generation time. This causes issues if the tool ever needs to be debugged, as this code generation will need to occur with each edit to any part of the program; additionally, this is lost time that could be spent testing and analyzing the code. This is one of the more prominent issues with the auto-coding process, and while much information is available with regard to optimizing Simulink designs to produce efficient and reliable C++ code, not much research has been made public on how to reduce the code generation time. It is of interest to develop some insight as to what causes code generation times to be so significant, and determine if there are architecture guidelines or a desirable auto-coding configuration set to assist in streamlining this step of the design process for particular applications. To address the issue at hand, the Simulink coder was studied at a foundational level. For each different component type made available by the software, the features, auto-code generation time, and the format of the generated code were analyzed and documented. Tools were developed and documented to expedite these studies, particularly in the area of automating sequential builds to ensure accurate data was obtained. Next, the Ramses model was examined in an attempt to determine the composition and the types of technologies used in the model. This enabled the development of a model that uses similar technologies, but takes a fraction of the time to auto-code to reduce the turnaround time for experimentation. Lastly, the model was used to run a wide array of experiments and collect data to obtain knowledge about where to search for bottlenecks in the Ramses model. The resulting contributions of the overall effort consist of an experimental model for further investigation into the subject, as well as several automation tools to assist in analyzing the model, and a reference document offering insight to the auto-coding process, including documentation of the tools used in the model analysis, data illustrating some potential problem areas in the auto-coding process, and recommendations on areas or practices in the current Ramses model that should be further investigated. Several skills were required to be built up over the course of the internship project. First and foremost, my Simulink skills have improved drastically, as much of my experience had been modeling electronic circuits as opposed to software models. Furthermore, I am now comfortable working with the Simulink Auto-coder, a tool I had never used until this summer; this tool also tested my critical thinking and C++ knowledge as I had to interpret the C++ code it was generating and attempt to understand how the Simulink model affected the generated code. I had come into the internship with a solid understanding of Matlab code, but had done very little in using it to automate tasks, particularly Simulink tasks; along the same lines, I had rarely used shell script to automate and interface with programs, which I gained a fair amount of experience with this summer, including how to use regular expression. Lastly, soft-skills are an area everyone can continuously improve on; having never worked with NASA engineers, which to me seem to be a completely different breed than what I am used to (commercial electronic engineers), I learned to utilize the wealth of knowledge present at JSC. I wish I had come into the internship knowing exactly how helpful everyone in my branch would be, as I would have picked up on this sooner. I hope that having gained such a strong foundation in Simulink over this summer will open the opportunity to return to work on this project, or potentially other opportunities within the division. The idea of leaving a project I devoted ten weeks to is a hard one to cope with, so having the chance to pick up where I left off sounds appealing; alternatively, I am interested to see if there are any opening in the future that would allow me to work on a project that is more in-line with my research in estimation algorithms. Regardless, this summer has been a milestone in my professional career, and I hope this has started a long-term relationship between JSC and myself. I really enjoy the thought of building on my experience here over future summers while I work to complete my PhD at Missouri University of Science and Technology.

Gualdoni, Matthew J.↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗