Search NASA⌕ Search

SEARCH · Search NASA

Results for “model data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Resolving temperature limitation on spring productivity in an evergreen conifer forest using a model–data fusion framework

Abstract. The flow of carbon through terrestrial ecosystems and the response to climate are critical but highly uncertain processes in the global carbon cycle. However, with a rapidly expanding array of in situ and satellite data, there is an opportunity to improve our mechanistic understanding of the carbon (C) cycle's response to land use and climate change. Uncertainty in temperature limitation on productivity poses a significant challenge to predicting the response of ecosystem carbon fluxes to a changing climate. Here we diagnose and quantitatively resolve environmental limitations on the growing-season onset of gross primary production (GPP) using nearly 2 decades of meteorological and C flux data (2000–2018) at a subalpine evergreen forest in Colorado, USA. We implement the CARbon DAta-MOdel fraMework (CARDAMOM) model–data fusion network to resolve the temperature sensitivity of spring GPP. To capture a GPP temperature limitation – a critical component of the integrated sensitivity of GPP to temperature – we introduced a cold-temperature scaling function in CARDAMOM to regulate photosynthetic productivity. We found that GPP was gradually inhibited at temperatures below 6.0 ∘C (±2.6 ∘C) and completely inhibited below −7.1 ∘C (±1.1 ∘C). The addition of this scaling factor improved the model's ability to replicate spring GPP at interannual and decadal timescales (r=0.88), relative to the nominal CARDAMOM configuration (r=0.47), and improved spring GPP model predictability outside of the data assimilation training period (r=0.88). While cold-temperature limitation has an important influence on spring GPP, it does not have a significant impact on integrated growing-season GPP, revealing that other environmental controls, such as precipitation, play a more important role in annual productivity. This study highlights growing-season onset temperature as a key limiting factor for spring growth in winter-dormant evergreen forests, which is critical in understanding future responses to climate change.

54 ENVIRONMENTAL SCIENCES↗

CORAL: A framework for rigorous self-validated data modeling and integrative, reproducible data analysis

Abstract Background Many organizations face challenges in managing and analyzing data, especially when relevant datasets arise from multiple sources and methods. Analyzing heterogeneous datasets and additional derived data requires rigorous tracking of their interrelationships and provenance. This task has long been a Grand Challenge of data science and has more recently been formalized in the FAIR principles: that all data objects be Findable, Accessible, Interoperable, and Reusable, both for machines and for people. Adherence to these principles is necessary for proper stewardship of information, for testing regulatory compliance, for measuring the efficiency of processes, and for facilitating reuse of data-analytical frameworks. Findings We present the Contextual Ontology-based Repository Analysis Library (CORAL), a platform that greatly facilitates adherence to all 4 of the FAIR principles, including the especially difficult challenge of making heterogeneous datasets Interoperable and Reusable across all parts of a large, long-lasting organization. To achieve this, CORAL's data model requires that data generators extensively document the context for all data, and our tools maintain that context throughout the entire analysis pipeline. CORAL also features a web interface for data generators to upload and explore data, as well as a Jupyter notebook interface for data analysts, both backed by a common API. Conclusions CORAL enables organizations to build FAIR data types on the fly as they are needed, avoiding the expense of bespoke data modeling. CORAL provides a uniquely powerful platform to enable integrative cross-dataset analyses, generating deeper insights than are possible using traditional analysis tools.

97 MATHEMATICS AND COMPUTING↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 2. Evaluating Controls on Flow Persistence in an Urbanized Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in an urbanized catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, distributed temperature sensing (DTS), continuous self-potential (SP) monitoring, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Field_Application subfolder contains the ATS XML input scripts, data files, output data for the SP site. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. The flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.m can only be used with COMSOL with MATLAB) is executed using the ATS output data to simulate the potential field. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) DTS Contains collated DTS data including raw Stokes and anti-Stokes measurement (provided as .h5 file). It also includes DTS processing.ipynb, a Jupyter notebook for calibrating the DTS data using dts_calibration Python package. cooler_calibration.csv is the DTS calibration CSV used in the calibration sequence. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion. 6) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 7) SP Contains the SP data collected in field at the SP sites (provided as CSV files). 8) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). Note: Code files (.ipynb, .py, .xml) can be opened in any standard code editor, .exo file can be viewed using Paraview, .h5 files can be opened using HDFView software and h5py Python package, and .resipy file can be opened with the open-source ResIPy software.

ATS↗

Using machine learning and artificial intelligence to improve model-data integrated earth system model predictions of water and carbon cycle extremes

The research proposed here focuses on improving the predictive power of the land component of earth system models (ESMs) using (1) model-data fusion enabled by machine learning (ML) and artificial intelligence (AI), (2) predictive modeling through the combination of ML, AI, and big-data (comprising both model output and observations), and (3) insight of ESM structure and process mechanisms gleaned from complex data using ML and AI.

54 ENVIRONMENTAL SCIENCES↗

Model, data, and code for paper "Modeling of streamflow in a 30-kilometer-long reach spanning 5 years using OpenFOAM 5.x"

The data package includes data, model, and code that support the analyses and conclusions in the paper titled “modeling of streamflow in a 30-kilometer-long reach spanning 5 years using OpenFOAM 5.x”. The primary goal of this paper is to demonstrate that key streamflow properties such as water depth, flow velocity, and dynamic pressure in a natural river at 30-kilometer scale over 5 years can be reliably and efficiently modeled using the computational framework presented in this paper. To support the paper, various data types from remote sensing, field observations, and computational models are used. Specific details are described as follows. Firstly, the river bathymetry data was obtained from a Light Detection and Ranging (LiDAR) survey. This data is then converted to a triangulated surface format, STL, for mesh generation in OpenFOAM. The STL data can be found in Model_Setups/BaseCase_2013To2015/constant/triSurface. The OpenFOAM mesh generated using this STL file can be found in constant/polyMesh. Other model setups, boundary and initial conditions can be found in /system and /0.org under folder BaseCase_2013To2015. A similar data structure can also be found in BaseCase_2018To2019 for the simulations during 2018 and 2019. Secondly, the OpenFOAM simulations need the upstream discharge and water depth information at the upstream boundary to drive the model. These data are generated from a one-dimensional hydraulic model and the data can be found under the folder Model_Setups /1D model Mass1 data. The mass1_65.csv and mass1_191.csv files include the results of the 1D model at the model inlet and outlet, respectively. The Matlab source code Mass1ToOFBC20182019.m is used to convert these data into OpenFOAM boundary condition setups.With the above OpenFOAM model, it can generate data for water surface elevation, flow velocity, and dynamic pressure. In this paper, the water surface elevation was measured at 7 locations during different periods between 2011 and 2019. The exact survey locations (see Fig1_SurveyLocations.txt) can be found in folder Fig_1. The variation of water stage over time at the 7 locations can be found in folder /Observation_WSE. The data type include .txt, .csv, .xlsx, and .mat. The .mat data can be loaded by Matlab.We also measured the flow velocities at 12 cross-sections along the river. At each cross-section, we recorded the x, y locations, depth, three velocity components u,v,w. These data are saved to a Matlab format which can be found under folder /Observation_Velocity and /Fig_1. The relative locations of velocity survey locations to the river bathymetry can be found in Figure 1c.The water stage data at the 7 locations from OpenFOAM, 1D, and 2D hydraulic models are also provided to evaluate the long-term performance of 3D models vs 1D/2D models. The water stage data for the 7 locations from OpenFOAM have been saved to .mat format and can be found in /OpenFOAM_WSE. The water stage data from the 1D model are saved in .csv format and can be found in /Mass1_WSE. The water stage from the 2D model is saved as .mat format and can be found in / Mass2_WSEIn addition, the OpenFOAM model outputs the information of hydrostatic and hydrodynamic pressure. They are saved as .mat format under folder /Fig_11/2013_1. As the files are too large, we only uploaded the data for January 2013. The area of different ratio of dynamic pressure to static pressure for all simulation range, i.e., 2013-2015, are saved to .mat format. They can be found in /Fig_11/PA. Further, the data of wall clock time versus the solution time of the OpenFOAM modeling are also saved to .mat format under folder /Fig_13/LogsMat. In summary, the data package contains seven data types, including .txt, .csv, .xlsx, .dat, .stl, .m, and .mat. The former 4 types can be directly open using a text editor or Microsoft Office. The .mat format needs to be read by Matlab. The Matlab source code .m files need to be run with Matlab. The OpenFOAM setups can be visualized in ParaView. The .stl file can be opened in ParaView or Blender. The data in subfolders Fig_1 to Fig_10 and Fig_12 are copied from the aforementioned data folders to generate specific figures for the paper. A readME.txt file is included in each subfolder to further describe how the data in each folder are generated and used to support the paper.Please use the data package's DOI to cite the data package. Please contact yunxiang.chen@pnnl.gov if you need more data related to the paper.

54 ENVIRONMENTAL SCIENCES↗

GDM (Grid Data Models) [SWR-24-31]

Grid Data Models (GDM) is a collection of pydantic data models for power distribution and transmission system components. These data models can be easily serialized and deserialized. GDM provides a standard interface for accessing power system component data with built in validation. See also: https://pypi.org/project/grid-data-models/

Duwadi, Kapil↗

Data and Scripts Associated with "Modeling Ecohydrological Responses of Vegetation to Urban Microclimates Using the E3SM Land Model"

This dataset supports the study of vegetation ecohydrological responses to urban microclimates using the land component of the Energy Exascale Earth System Model (ELM) at four urban sites in Knoxville, Tennessee, USA. It includes the model inputs, simulation outputs, and associated scripts for running ELM simulations and analyzing the resulting data. The Model_Inputs folder includes static surface data, satellite-derived phenology (i.e., leaf area index), and atmospheric forcing data used to drive ELM simulations. Detailed descriptions of these datasets are provided in Section 2.3.2 of the associated manuscript. The Model_Outputs folder contains simulation results for the baseline, treatment, and ensemble experiments. Outputs from the baseline and treatment simulations are provided as raw ELM NetCDF files. Because the raw outputs from the 4,000-member ensemble are prohibitively large, the ensemble results are provided as summarized CSV files, which also serve as the source data for Figure 5 of the associated manuscript. The Scripts folder contains three components: E3SM, the core codebase of the Energy Exascale Earth System Model (E3SM); elm-olmt, the Offline Land Model Testbed (OLMT) used to perform the simulations; and knoxville_elm, which contains the analysis scripts used to process model outputs and generate the figures and results presented in the associated manuscript. Additional information is provided in Scripts_readme.txt within the Scripts directory.

Lu, Xiaoman [ORNL] (ORCID:0000000306698780)↗

ESS-DIVE guidelines for archiving terrestrial model data

This dataset contains supporting documents and images for ESS-DIVE terrestrial model data archiving guidelines.Terrestrial models are broadly defined as numerical models that couple both land dynamics and energy, water, carbon, or nutrient fluxes. We created these guidelines based on input from the U.S. Department of Energy’s Biological and Environmental Research land modeling community. The guidelines are intended to help modelers determine which components of their terrestrial model data associated with publication should be archived. Based on input from the land modeling community, the guidelines recommend archiving both model input and testing data, as well as code, script, and metadata. The guidelines also recommend archiving model data output, depending on the limitations set by data repositories. Lastly, we provide recommendations for bundling data files for publication as well as a discussion about tools that can facilitate model data archiving and reuse.This dataset is an archive of the associated GitHub repository for our model archiving guidelines (https://github.com/ess-dive-community/essdive-model-data-archiving-guidelines). The ‘README.pdf’ file gives a general introduction to the guidelines, and the ‘instructions.pdf’ file provides more detailed steps for following the guidelines. We also provide 2 figures in this data package: 1) a decision tree (model_data_guidelines_decision_tree.png) that can help users determine which components of their model data to archive. and 2) the ‘model_data_guidelines_flmd.png’ file depicts the different files that can be archived in addition to the model data itself. Lastly, we include 3 digitized tables from our associated manuscript and 3 CSV files with anonymized input from DOE scientists about the importance of different aspects of model data archiving from which we developed the guidelines.Dataset updates for v1.1.0: We updated this data package on 2021-11-22 in response to review comments on our related manuscript. In this update we removed one figure so that the model archiving guidelines are conveyed in text rather than an image. We updated the file-level metadata (FLMD) figure to be in accord with the most recent FLMD recommendations. We made minor edits to the README file to update the recommended citation and added two co-authors. We also added 6 new data files (3 are anonymized input from DOE scientists that helped to inform guidelines, and 3 are digitized tables from our manuscript.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning for Gearbox Fault Prediction by Using Both Scada and Modeled Data

This presentation outlines the work in the paper titled "Prognostics of Wind Turbine Gearbox Bearing Failures Using SCADA and Modeled Data" published by the PHM Society and presented at its 2020 annual conference. It is accessible at https://papers.phmsociety.org/index.php/phmconf/article/download/1292/862. The technical work is on machine learning approaches for prognostics for gearbox faults. The methodology combines SCADA time series data and physics domain modeling data, derived from the models developed by the NREL team, as inputs to machine learning models to predict gearbox bearing failures with one month lead time. Based on SCADA data, modeled data, and bearing failure log data from an actual wind plant, the performances of different machine-learning models on unseen data are then evaluated using industry-standard metrics such as precision, recall, and F1 score, and AUC (area under receiver operating characteristic curve). Results show the overall system performance enhancement in predicting bearing failure when modeled data are included with SCADA data. The reduction in terms of false alarms is about 50%, and improvement in terms of precision, and F1 score, and AUC is about 33%, and 12%, and 6% respectively, based on the best modeling case in this study.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

Small Scale WEC Performance Modeling Data

Small Scale WEC Performance Modeling Data is performance data from downscaled models of common WEC devices and their calculated performance outputs. This data is used by the Small WEC interactive modeling tool hosted by PRIMRE. The devices include a point absorber, a two-body point absorber (RM3), an oscillating surge device (OSWEC), and an attenuator type device (McCabe Wave Pump). One of the primary use cases for this work is to give an easy way to compare power output for a variety of WECs and model sizes.

16 TIDAL AND WAVE POWER↗

Prognosis of Wind Turbine Gearbox Bearing Failures Using SCADA and Modeled Data

Predictive maintenance and condition monitoring systems for wind turbines have seen increased adoption to minimize downtime, reducing operation and maintenance costs. On today’s wind power plants, the integrated supervisory control and data acquisition (SCADA) system provides low- frequency operational data that can be leveraged to quantify a wind turbine’s health. The aim of this study is to utilize machine-learning techniques to predict axial cracking failures in wind turbine gearbox bearings up to 1 month ahead of time. The failures are assumed to have occurred when the investigated bearing was replaced. While current SCADA systems show the overall condition of a wind turbine, often they do not allow for the investigation of specific gearbox bearings’ health. To enrich bearing fault signatures, additional data are computed through physics-based models using gearbox design information. Based on SCADA data, modeled data, and bearing failure log data from an actual wind plant, the performances of different machine-learning models on unseen data are then evaluated using industry-standard metrics such as precision, recall, and F1 score. Results show the overall system performance enhancement in predicting bearing failure when modeled data are included with SCADA data. The reduction in terms of false alarms is about 50%, and improvement in terms of precision and F1 score is about 33% and 12% respectively, based on the best modeling case in this study.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

A generalized approach to x-ray data modeling for high-energy-density plasma experiments

Accurate understanding of x-ray diagnostics is crucial for both interpreting high-energy-density experiments and testing simulations through quantitative comparisons. X-ray diagnostic models are complex. Past treatments of individual x-ray diagnostics on a case-by-case basis have hindered universal diagnostic understanding. Here, in this study, we derive a general formula for modeling the absolute response of non-focusing x-ray diagnostics, such as x-ray imagers, one-dimensional space-resolved spectrometers, and x-ray power diagnostics. The present model is useful for both data modeling and data processing. It naturally accounts for the x-ray crystal broadening. The new model verifies that standard approaches for a crystal response can be good approximations, but they can underestimate the total reflectivity and overestimate spectral resolving power by more than a factor of 2 in some cases near reflectivity edge features. We also find that a frequently used, simplified-crystal-response approximation for processing spectral data can introduce an absolute error of more than an order of magnitude and the relative spectral radiance error of a factor of 3. The present model is derived with straightforward geometric arguments. It is more general and is recommended for developing a unified picture and providing consistent treatment over multiple x-ray diagnostics. Such consistency is crucial for reliable multi-objective data analyses.

Nagayama, Taisuke↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗