Search NASA⌕ Search

SEARCH · Search NASA

Results for “predictability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Initialized Earth system prediction from subseasonal to decadal timescales

Initialized Earth system predictions are made by starting a numerical prediction model in a state as consistent as possible to observations, and running it forward in time for up to ten years. Skillful predictions at time slices from subseasonal to seasonal (S2S), seasonal to interannual (S2I) and seasonal to decadal (S2D) offer information useful for various stakeholders, from agriculture to water resource management, and human and infrastructure safety. In this Review, we examine the processes influencing predictability, and discuss estimates of skill across S2S, S2I and S2D timescales. There are encouraging signs that skillful predictions can be made: at S2S timescales, there has been some skill in predicting the Madden-Julian Oscillation and North Atlantic Oscillation; at S2I in predicting the El Niño-Southern Oscillation; and at S2D, in predicting variability in North Atlantic sea surface temperatures. However, challenges remain, and future work must prioritise reducing model error, more effectively communicating forecasts to users, and increasing process and mechanistic understanding that could increase predictive skill and, in turn, confidence. As numerical models progress towards Earth system models, initialized predictions are expanding to include prediction of sea-ice, air pollution, terrestrial and ocean biochemistry which can bring clear benefit to society and various stakeholders.

climate prediction↗

Improving protein tertiary structure prediction by deep learning and distance prediction in CASP14

Abstract Substantial progresses in protein structure prediction have been made by utilizing deep‐learning and residue‐residue distance prediction since CASP13. Inspired by the advances, we improve our CASP14 MULTICOM protein structure prediction system by incorporating three new components: (a) a new deep learning‐based protein inter‐residue distance predictor to improve template‐free (ab initio) tertiary structure prediction, (b) an enhanced template‐based tertiary structure prediction method, and (c) distance‐based model quality assessment methods empowered by deep learning. In the 2020 CASP14 experiment, MULTICOM predictor was ranked seventh out of 146 predictors in tertiary structure prediction and ranked third out of 136 predictors in inter‐domain structure prediction. The results demonstrate that the template‐free modeling based on deep learning and residue‐residue distance prediction can predict the correct topology for almost all template‐based modeling targets and a majority of hard targets (template‐free targets or targets whose templates cannot be recognized), which is a significant improvement over the CASP13 MULTICOM predictor. Moreover, the template‐free modeling performs better than the template‐based modeling on not only hard targets but also the targets that have homologous templates. The performance of the template‐free modeling largely depends on the accuracy of distance prediction closely related to the quality of multiple sequence alignments. The structural model quality assessment works well on targets for which enough good models can be predicted, but it may perform poorly when only a few good models are predicted for a hard target and the distribution of model quality scores is highly skewed. MULTICOM is available at https://github.com/jianlin-cheng/MULTICOM_Human_CASP14/tree/CASP14_DeepRank3 and https://github.com/multicom-toolbox/multicom/tree/multicom_v2.0 .

59 BASIC BIOLOGICAL SCIENCES↗

MULTICOM2 open-source protein structure prediction system powered by deep learning and distance prediction

Protein structure prediction is an important problem in bioinformatics and has been studied for decades. However, there are still few open-source comprehensive protein structure prediction packages publicly available in the field. In this paper, we present our latest open-source protein tertiary structure prediction system—MULTICOM2, an integration of template-based modeling (TBM) and template-free modeling (FM) methods. The template-based modeling uses sequence alignment tools with deep multiple sequence alignments to search for structural templates, which are much faster and more accurate than MULTICOM1. The template-free (ab initio or de novo) modeling uses the inter-residue distances predicted by DeepDist to reconstruct tertiary structure models without using any known structure as template. In the blind CASP14 experiment, the average TM-score of the models predicted by our server predictor based on the MULTICOM2 system is 0.720 for 58 TBM (regular) domains and 0.514 for 38 FM and FM/TBM (hard) domains, indicating that MULTICOM2 is capable of predicting good tertiary structures across the board. It can predict the correct fold for 76 CASP14 domains (95% regular domains and 55% hard domains) if only one prediction is made for a domain. The success rate is increased to 3% for both regular and hard domains if five predictions are made per domain. Moreover, the prediction accuracy of the pure template-free structure modeling method on both TBM and FM targets is very close to the combination of template-based and template-free modeling methods. This demonstrates that the distance-based template-free modeling method powered by deep learning can largely replace the traditional template-based modeling method even on TBM targets that TBM methods used to dominate and therefore provides a uniform structure modeling approach to any protein. Finally, on the 38 CASP14 FM and FM/TBM hard domains, MULTICOM2 server predictors (MULTICOM-HYBRID, MULTICOM-DEEP, MULTICOM-DIST) were ranked among the top 20 automated server predictors in the CASP14 experiment. After combining multiple predictors from the same research group as one entry, MULTICOM-HYBRID was ranked no. 5. The source code of MULTICOM2 is freely available at https://github.com/multicom-toolbox/multicom/tree/multicom_v2.0 .

97 MATHEMATICS AND COMPUTING↗

Prediction of histone post-translational modifications using deep learning

Abstract Motivation Histone post-translational modifications (PTMs) are involved in a variety of essential regulatory processes in the cell, including transcription control. Recent studies have shown that histone PTMs can be accurately predicted from the knowledge of transcription factor binding or DNase hypersensitivity data. Similarly, it has been shown that one can predict PTMs from the underlying DNA primary sequence. Results In this study, we introduce a deep learning architecture called DeepPTM for predicting histone PTMs from transcription factor binding data and the primary DNA sequence. Extensive experimental results show that our deep learning model outperforms the prediction accuracy of the model proposed in Benveniste et al. (PNAS 2014) and DeepHistone (BMC Genomics 2019). The competitive advantage of our framework lies in the synergistic use of deep learning combined with an effective pre-processing step. Our classification framework has also enabled the discovery that the knowledge of a small subset of transcription factors (which are histone-PTM and cell-type-specific) can provide almost the same prediction accuracy that can be obtained using all the transcription factors data. Availabilityand implementation https://github.com/dDipankar/DeepPTM. Supplementary information Supplementary data are available at Bioinformatics online.

Baisya, Dipankar Ranjan (ORCID:0000000267847359)↗

Predicting nepheline precipitation in waste glasses using ternary submixture model and machine learning

Nepheline precipitation in nuclear waste glasses during vitrification can be detrimental due to its negative effect on chemical durability. Developing models to accurately predict nepheline precipitation from compositions is important to increase waste loading since existing models can be overly conservative. In this study, an expanded dataset containing 955 glasses was compiled from literature data, where 355 glasses are for high-level waste (HLW). Previously developed submixture models were refitted using the new dataset, where a misclassification rate of 7.8% was achieved. Nine machine learning (ML) algorithms (e.g., k-nearest neighbor, Gaussian process regression, artificial neural network, support vector machine, decision tree, etc.) were applied to evaluate their ability of predicting nepheline precipitation from compositions. Model accuracy, precision, recall/sensitivity, and F1 score were systemically compared between different ML algorithms and modeling protocols. Good model prediction with an accuracy ~0.9 (misclassification rate of ~10%) was observed with different algorithms under certain protocol. This study evaluated various ML models to predict nepheline precipitations in waste glasses, highlighting the importance of data preparation, modeling protocol, and their effect on model stability and reproducibility. The results provide insights into applying ML to predict glass properties and suggest areas for future research on modeling nepheline precipitations.

Lu, Xiaonan↗

Simulation-based Performance Evaluation of Model Predictive Control for Building Energy Systems

The performance of model predictive control (MPC) can be significantly affected by different choices of controller parameters such as the time intervals for model discretization and control sampling. Due to the lack of a systematic understanding on how these parameters affect control performance, they are usually selected arbitrarily in practice.In this paper, the combined impacts of selected time intervals for model discretization and control sampling on the performance of MPC are comprehensively investigated for the first time through detailed simulations. Specifically, a typical MPC strategy is first designed to improve building operations based on a reduced-order model of building dynamics. Then, the performance of the designed MPC is evaluated against different choices of time intervals for model discretization and control sampling on a simulated office building. The detailed simulation results reveal that the time interval for model discretization has a much greater influence on the performance of MPC than the time interval for control sampling. Although the time interval for control sampling usually receives more attentions in practice, it turns out that the time interval for model discretization affects the prediction performance, cost saving, and computation time simultaneously and more significantly. Therefore, the simulation-based performance evaluation presented here sheds light on the impacts of different time intervals and facilitates their selection for practical applications of MPC to building operations

Huang, Sen↗

Development of Digital Twin Predictive Model for PWR Components: Updates on Multi Times Series Temperature Prediction Using Recurrent Neural Network, DMW Fatigue Tests, System Level Thermal-Mechanical-Stress Analysis

The long-term operation (LTO) of nuclear power plant (NPP) beyond their original design life of 40 years, can lead to more material damage associated with cyclic fatigue under thermal-mechanical loading cycles and associated long-term exposure of reactor material to the deleterious reactor-coolant environments. However, under this LTO condition the reactor components can still safely operate but may require more frequent Nondestructive Evaluation (NDE) of reactor components. Frequent NDE requirement may lead to frequent shutdown of the NPP. This in turn can lead to power outage and additional NDE-inspection-cost related economic loss. The economic loss can be minimized by reducing uncertainty in life estimation of safety-critical pressure boundary components and by implementing more digital approach such as by using upcoming digital-twin (DT) technology for predicting the structural states (e.g., time and location dependent inside/outside thickness temperature, stress, strain, plastic deformation, etc.) and associated fatigue life of a component in real time. Towards this goal Argonne National Laboratory (ANL) with the sponsorship of DOE Light Water Reactor Sustainability (LWRS) program is working on the development of a DT framework that can be used for real time environmental fatigue prediction of reactor components. The DT framework is based on limited experiment-data, Artificial-intelligence (AI) – Machine-Learning (ML) - Deep-Learning (DL) based techniques and Multiphysics-computational-mechanics such as finite element (FE) based modeling tools. Towards this overall goal, following are some of the major contributions made during the FY21: 1) Multiple 82/182 dissimilar metal weld (DMW) specimens (both solid-weld and joint-weld representing the actual reactor multi-metal nozzles) were fatigue tested. The resulting fatigue lives were compared to the NUREG-6909 based best-fit and design fatigue curves. Additionally, the results of 52/152 DMW fatigue specimens (which were recently tested at Republic of Korea under the sponsorship of International Nuclear Energy Research Initiative - INERI program) were compared to the NUREG-6909 based best-fit and design fatigue curves. From the comparison of 82/182 and 52/152 DMW test data with NUREG-6909 best-fit curve, most of the reported test data fall way away from the NUREG-6909 suggested best-fit or mean curve. The NUREG-6909 suggested best-fit curve is the best-fit curve of austenitic stainless steel and due to lack of enough data on Nickel-based welds, this is currently being used for predicting the life of Nickel-alloy-based welded components. However, the above observation may require higher scaling factor (e.g., ASME suggested factor of 20 on cycles rather than the current NUREG-6909 suggested factor of 12 on cycles) for scaling the austenitic-stainless-steel best-fit-curve for estimating the design or safe-life of a welded component. Accordingly, for example, if a DMW component experience a strain amplitude of 0.6% the PWR-water life of the component would be 52 cycles instead of 85 cycles. However, more DMW tests are required to further ascertain the above-mentioned observations. 2) A system level CAD and finite element model were developed which consists of reactor pressure vessel (RPV), part of steam generator (SG), part of pressurizer (PRZ), hot leg (HL), and surge line (SL). This is with detailed nozzle geometry and thermal-mechanical material properties of different metals to simulate realistic thermal-mechanical stress under connected system global thermal-mechanical boundary conditions. 3) Different system level heat transfer analyses were performed with estimation of relevant heat transfer coefficients. The resulting data were used in subsequent system level thermal-mechanical stress analysis and for generating spatial-temporal training and validation data for a system level digital-twin based temperature predictor. Transient heat transfer analyses were performed considering thermal boundary condition under design-basis (DB) loading and EDF (Électricité de France) data-based grid-load-following (EDF-GLF) loading cycles. 4) System level thermal-mechanical stress analysis was performed for identifying damage-prone hotspots and for future extension of the model for cyclic state prediction. From the system-level model simulation under DB loading cycle it is found that HL and the SL nozzle that connect to the HL can experience significant stress and strain and could be one of the weakest links in the overall reactor coolant system (RCS). 5) An AI/ML based DT model was developed for multi-time-series temperature prediction at any inside/outside thickness locations of PWR pressure boundary components. This is by using Recurrent-neural-network (RNN) and keras machine learning libraries. The RNN model was validated against two laboratory test-based data sets with one obtained through ANL’s in-air fatigue test system and other through PWR-water test loop. The experimentally validated DT model further validated against FE model results to predict thermal scarification related spatialtemporal temperatures at random locations of a component. The well validated DT model was then used for demonstrating spatial-temporal temperature prediction under 100+ years of reactor operation subjected to combined DB, EDF-GLF and randomized grid-load-following (RANDOMGLF) loading Cycles. The expert-elicitation DT model framework was developed assuming field/input/process measurements can be available from a few existing plant sensors and can readily be used by the NPP operators. The above temperature prediction model will feed to the next-step stress analysis model based on which the life of a component can be predicted in realtime, which is one of our future works.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

DeepComplex: A Web Server of Predicting Protein Complex Structures by Deep Learning Inter-chain Contact Prediction and Distance-Based Modelling

Proteins interact to form complexes. Predicting the quaternary structure of protein complexes is useful for protein function analysis, protein engineering, and drug design. However, few user-friendly tools leveraging the latest deep learning technology for inter-chain contact prediction and the distance-based modelling to predict protein quaternary structures are available. To address this gap, we develop DeepComplex, a web server for predicting structures of dimeric protein complexes. It uses deep learning to predict inter-chain contacts in a homodimer or heterodimer. The predicted contacts are then used to construct a quaternary structure of the dimer by the distance-based modelling, which can be interactively viewed and analysed. The web server is freely accessible and requires no registration. It can be easily used by providing a job name and an email address along with the tertiary structure for one chain of a homodimer or two chains of a heterodimer. The output webpage provides the multiple sequence alignment, predicted inter-chain residue-residue contact map, and predicted quaternary structure of the dimer.

59 BASIC BIOLOGICAL SCIENCES↗

Prediction of performance and turbulence in ITER burning plasmas via nonlinear gyrokinetic profile prediction

Burning plasma performance, transport, and the effect of hydrogen isotope (H, D, D-T fuel mix) on confinement has been predicted for ITER baseline scenario (IBS) conditions using nonlinear gyrokinetic profile predictions. Accelerated by surrogate modeling (Rodriguez-Fernandez et al 2022 Nucl. Fusion 62 076036), high fidelity, nonlinear gyrokinetic simulations performed with the CGYRO code (Candy et al 2016 J. Comput. Phys. 324 73), were used to predict profiles of T i , T e , and n e while including the effects of alpha heating, auxiliary power (NBI + ECH), collisional energy exchange, and radiation losses inside of $r/a$ = 0.9. Predicted profiles and resulting energy confinement are found to produce fusion power and gain that are approximately consistent with mission goals ($P_\textrm{fusion} = 500$ MW at Q = 10) for the baseline scenario and exhibit energy confinement that is within 1σ of the H-mode energy confinement scaling. The power of the surrogate modeling technique is demonstrated through the prediction of alternative ITER scenarios with reduced computational cost. These scenarios include conditions with maximized fusion gain and an investigation of potential resonant magnetic perturbation (RMP) effects on performance with a minimal number of gyrokinetic profile iterations required (3–6). These predictions highlight the stiff ITG nature of the core turbulence predicted in the ITER baseline and demonstrate that $Q \gt$ 17 conditions may be accessible by reducing auxiliary input power while operating in IBS conditions. Prediction of full kinetic profiles allowed for the projection of hydrogen isotope effects around ITER baseline conditions. The gyrokinetic fuel ion species was varied from H, D, and 50/50 D-T and kinetic profiles were predicted. Results indicate that a weak or negligible isotope effect will be observed to arise from core turbulence in IBS conditions. The resulting energy confinement, turbulence, and density peaking, and the implications for ITER operations will be discussed.

gyrokinetics↗

Hybridizing Machine Learning and Physically-based Earth System Models to Improve Prediction of Multivariate Extreme Events (AI Exploration of Wildland Fire Prediction)

Focal Areas: This project responds to two focal areas identified in the DOE Call for AI4ESP White Papers: 1) Predictive modeling through the use of artificial intelligence (AI) techniques, and 2) insights gleaned from complex data using explainable AI and big data analytics. Science Challenge: Large wildland fires (hereafter wildfires) appearing as high-impact compound climate extreme events are closely related to hydroclimate and water cycle extremes that modulate surface fuel supply and combustibility. These compound events have multivariate climatic features (e.g., temperature, precipitation, relative humidity, wind, lightning) and societal drivers (e.g., forest management, land use change, human caused ignitions). Meanwhile, they induce strong feedbacks to the coupled atmosphere, biosphere, and hydrosphere by perturbing regional and global radiation budget as well as ecological, biogeochemical, and water cycles across multiple spatiotemporal scales. The nonlinear interactions between these natural and anthropogenic components of the Earth system are too complex to be completely and adequately represented in today’s Earth system models (ESMs). The inherent stochastic nature of fire activity at all scales further increases the difficulty of its prediction using ESMs that are usually developed from deterministic equations and parameterizations. Besides, concurrence of long-term (decadal to interdecadal) global climate change and fire regime shifts overlapping with short-term (intraseasonal to interannual) variations of regional fire weather and burning activity confound predictability of these compound extreme events. We propose to address the above scientific challenges by using machine learning (ML)-based data-driven modeling techniques to integrate observations and physically-based ESMs’ simulations in a computationally efficient hybrid prediction system. This prediction system is supposed to characterize the wildfire’s sensitivity to climate and exogenous drivers at high resolution (~ 0.25°) on subseasonal to seasonal (S2S) timescales providing improved predictability and explainability. We will use the system to help identify: (1) What are the computational elements of a hybrid system needed to predict compound climate extreme events such as global wildfires? (2) What are the key drivers (either natural or anthropogenic) that modulate short-term variations of multivariate fire weather and burning activity over different regions? How can one take advantage of those driver-response relationships to improve the predictability of large wildfires on S2S time scales? (3) What are the underlying physical mechanisms and sources of improved predictability? Which ML techniques are optimal in revealing and adapting these mechanisms?

54 ENVIRONMENTAL SCIENCES↗

The Solar Influencer Next Door: Predicting Low-Income Solar Referrals and Leads

Increasing the adoption of solar among low-to-moderate income (LMI) households remains an important policy goal because of its promise to simultaneously reduce energy burden and support the just distribution of benefits of renewable energy. However, scaling LMI solar remains challenging due to affordability and access issues. Most existing LMI adoption has occurred under public-funded programs, highlighting the importance of increasing the cost-effectiveness of these programs at scale. We develop a new household-level data set on LMI solar lead acquisition, referrals, and adoption to understand the processes through which LMI solar uptake has occurred in California. Then, we develop models to predict two sub-mechanisms in the solar adoption process: whether an otherwise qualified lead becomes "lost" i.e. non-responsive to outreach and, for existing clients, whether they refer solar to others. For the program analyzed, participants received their solar system at no cost, which deemphasizes economic drivers of solar adoption and could differ from other program experiences. Both models substantially improved the accuracy of prediction relative to a baseline. Overall, we find that peer effects and solar economics are important to predicting referrals, and household demographic factors in lead loss prediction. Finally, we find that referrals are both the highest quality and largest source of LMI solar leads, providing a promising mechanism to expand LMI programs further.

customer acquisition costs↗

A review of machine learning in building load prediction

The surge of machine learning in recent years has been empowering engineer modeling in various fields. The decreasing hardware cost, increasing data accessibility, and advances of building automation system (BAS) allow the collection and storage of a significant amount of building operation data. The two facts provide great opportunities of applying machine learning to building energy systems modeling and analysis. There are a great number of research papers on this topic but there lacks a comprehensive and general review to summarize the current development, limitations, gaps and future trend. In this review paper series, machine learning techniques in building energy system modeling and analysis are reviewed under the organization and logic of the machine learning definition by Tom M. Mitchell: a computer program is said to learn from experience E with respect to some class of tasks T and performance measure P if its performance at tasks in T, as measured by P, improves with experience E. This paper is the first part of the review paper series, which focuses on building load prediction. First, the applications of building load prediction model (task T) are reviewed. Then, the modeling algorithms improving machine learning performance and accuracy (performance P) are reviewed. At the same time, the literature on the data perspective for modeling (experience E), including data engineering from sensors level to data level, pre-processing, feature extraction and selection, is reviewed. Finally, what is well-studied and what is lacking but with great potential are concluded; the gaps between present and future utilization of machine learning techniques are identified; the future trend and development are also predicted. The target readers of this paper are not only researchers from the building side who can get exposed to cutting edge machine learning tools, but also those from machine learning side who can understand the potential and challenge to apply machine learning in buildings.

Liang, Zhang↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

A machine learning approach for clinker quality prediction and nonlinear model predictive control design for a rotary cement kiln

Abstract Cement manufacturing is energy‐intensive (5Gj/t) and comprises a significant portion of the energy footprint of concrete systems. Incorporating modern monitoring, simulation and control systems will allow lower energy use, lower environmental impact, and lower costs of this widely used construction material. One of the goals of the CESMII roadmap project on the Smart Manufacturing of Cement included developing an analytical process model for clinker quality that includes the chemistry of the kiln feed and accounts for critical process variables. This predictive model will be used in nonlinear model predictive control system designed to significantly reduce process energy use while maintaining or improving product quality. In the cement manufacturing plant used in this study, the kiln feed (meal) is tested every 12 h and used to estimate the mineral composition of the cement kiln output (clinker) using the stoichiometry‐based Bogue's model and the expertise of the plant operators. During kiln operation, kiln output (clinker) is sampled and tested every 2 h to measure its chemical and mineral composition. The predicted and measured values of the clinker composition are used by the plant operators to adjust the kiln input stream and the production process characteristics to maintain stable operation and uniform product quality. However, the time delay between prediction and testing, along with inaccuracies inherent in the Bogue's model have made any process changes designed to minimize energy use problematic, especially in‐light of potential clinker quality issues that process changes often pose. A new analytical model that integrates quality information and process operation information has been developed from data collected from 2 years of production from an operating cement facility. To make the model fuel‐type‐independent, consumed heat energy was computed in the model instead of fuel type and amount. A Feedforward Network was trained and tailored from collected data. Many data‐based simulations were conducted to quantitatively evaluate the proposed model and the 5‐fold cross‐validation procedure was used to test the models. The resulting predictive model was shown to have a low root mean square error (MSE) with respect to the estimated clinker mineral composition compared to that using the industry standard “Bogue’ model”. The end goal of this work was to develop a single machine learning tool that allows the use of quality control data and process control variables to improve energy efficiency of the process in a continuous fashion. The proposed nonlinear model predictive control system (NMPC) can generate predicted kiln production characteristics based on manipulated variables in manner that accurately follows the target product quality values. Simulation results also show that the proposed model produced accurate predictions of kiln outputs that fell within the required constraints, while manipulating control variables within typical operational ranges.

Ali, Asem M.↗

Comparative Study on the Machine Learning-Based Prediction of Adsorption Energies for Ring and Chain Species on Metal Catalyst Surfaces

Computation of adsorption and transition state energies for a large number of surface intermediates for numerous active site models pose significant computational overhead in computational screening of catalysts. Machine learning (ML) techniques can be used to predict part of these energies. To predict the energies, ML models need to be fed appropriate metal and species descriptors. For complex surface chemistries, the structures of the intermediate species can vary greatly. In this paper, working with the hydrodeoxygenation of succinic acid on six different metal surfaces, we have studied the effect of linear and non-linear ML models used along with pen-and-paper based species descriptors and two categories of metal descriptors on two different categories of intermediate species: chain and ring. More specifically, our computations include the prediction of chain species when trained on only chain species and also when trained on both chain and ring species. Similar computations were performed for predictions of ring species. In each case, results of linear ML models were compared with kernel based non-linear models. Our results indicate that ring species data does not improve the prediction of chain species. Similarly, chain species data does not improve the prediction of ring species. The use of non-linear ML models, however, did help to minimize the prediction errors compared to the linear models. Furthermore, the study also shows that electronic or adsorption energy based metal descriptors along with bond count based species fingerprints can achieve a mean absolute error (MAE) of less than 0.2 eV for complex chain molecules when used with an appropriate machine learning model.

Adsorption↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗