Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Architecture & Analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Uncertainty quantification for molecular property predictions with graph neural architecture search

Graph Neural Networks (GNNs) have emerged as a prominent class of data-driven methods for molecular property prediction. However, a key limitation of typical GNN models is their inability to quantify uncertainties in the predictions. This capability is crucial for ensuring the trustworthy use and deployment of models in downstream tasks. To that end, we introduce AutoGNNUQ, an automated uncertainty quantification (UQ) approach for molecular property prediction. AutoGNNUQ leverages architecture search to generate an ensemble of high-performing GNNs, enabling the estimation of predictive uncertainties. Our approach employs variance decomposition to separate data (aleatoric) and model (epistemic) uncertainties, providing valuable insights for reducing them. In our computational experiments, we demonstrate that AutoGNNUQ outperforms existing UQ methods in terms of both prediction accuracy and UQ performance on multiple benchmark datasets, and generalizes well to out-of-distribution datasets. Additionally, we utilize t-SNE visualization to explore correlations between molecular features and uncertainty, offering insight for dataset improvement. AutoGNNUQ has broad applicability in domains such as drug discovery and materials science, where accurate uncertainty quantification is crucial for decision-making.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SENTRA: A Modular Computational Graph Framework for Critical Mineral and Materials Supply Chains: Part I: Network Construction Latent-Quantity Estimation, and Temporal Graph Forecasting

Global supply chains for critical minerals and materials are complex, evolving networks of countries, products, production stages, and trade relationships. Existing analytical approaches are limited by fragmented data and static network representations that do not capture the dynamic production dependencies linking raw materials, intermediate products, and final goods across multiple countries. Trade and production statistics provide only a partial view of domestic production, inventories, and material flows, making it difficult to identify indirect sourcing pathways, hidden dependencies, and embedded foreign exposures. This paper introduces the Supply Chain Exposure Network Tracking and Risk Assessment (SENTRA) framework, a modular graph-based computational framework for constructing, analyzing, and forecasting dynamic supply chain networks. As the first paper in a three-part methodological series, it establishes the computational foundation of SENTRA by constructing a temporal attributed multi-relational graph whose nodes represent product–country pairs and whose edges encode observed trade and within-country value-chain relationships. Statistical estimation and constrained optimization recover latent production, final demand, and product input dependency coefficients while enforcing economic accounting constraints. Graph-derived exposure measures quantify direct, transshipment, value-chain, and multi-hop supply chain dependencies independently of the forecasting model. A temporal graph forecasting architecture based on a relational graph neural network then forecasts the evolution of the graph under mass-balance constraints with distribution-free conformal uncertainty quantification. Validation on the global aluminum supply chain shows that the learned graph representations recover economically meaningful supply chain structure, accurately forecast out-of-sample trade relationships, and produce well-calibrated prediction intervals. Subsequent papers apply this computational foundation to exposure assessment, disruption analysis, and scenario-based policy analysis, and extend the framework to multimaterial supply chain modeling and decision support.

36 MATERIALS SCIENCE↗

A Comprehensive Chemistry Evaluation and Diagnostics Package for E3SM – ChemDyg Version 1.1.0

The Chemistry Evaluation and Diagnostics Package (ChemDyg) is an open-source tool designed for the Energy Exascale Earth System Model (E3SM) developed by the U.S. Department of Energy. ChemDyg facilitates routine evaluation, tailored development, and in-depth analysis of atmospheric chemistry through its modular architecture, allowing users to compare model outputs with observational data. Version 1.1.0 introduces a robust set of diagnostic capabilities, including climatology, time evolution of key tracers, diurnal and annual cycle analyses, and extensive budget diagnostics. These features help identify model discrepancies and enhance the representation of atmospheric chemistry in E3SM. Each self-contained diagnostic set includes dedicated scripts and documentation for ease of use. The interactive HTML output improves data accessibility, accelerating chemistry model development. Additionally, ChemDyg's flexible framework allows for customization, enabling users to create unique diagnostic sets for specific scientific contributions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Transferable Water Potentials Using Equivariant Neural Networks

Machine learning interatomic potentials (MLIPs) have emerged as a technique that promises quantum theory accuracy for reduced cost. It has been proposed [J. Chem. Phys. 2023, 158, 084111] that MLIPs trained on solely liquid water data cannot accurately transfer to the vapor–liquid equilibrium while recovering the many-body decomposition (MBD) analysis of gas-phase water clusters. This suggests that MLIPs do not directly learn the physically correct interactions of water molecules, limiting transferability. In this work, we show that MLIPs using equivariant architecture and trained on 3200 liquid water structures reproduces liquid-phase water properties (e.g., density within 0.003 g/cm 3 between 230 and 365 K), vapor–liquid equilibrium properties up to 550 K, the MBD analysis of gas-phase water cluster up to six-body interactions, and the relative energy and the vibrational density of states of ice phases. We show that potentials developed using equivariant MLIPs allow transferability for arbitrary phases of water that remain stable in nanosecond long simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrating Analytical Solutions and U-Net Model for Predicting Groundwater Contaminant Plumes in Pump-and-Treat Systems

Pump-and-treat (P&T) is a common technique for groundwater remediation involving the extraction and treatment of contaminated water above ground. Optimizing the design and operation of the P&T well network is essential for maximizing the system’s effectiveness and efficiency. However, this optimization often necessitates many model evaluations, leading to computationally demanding tasks. This study introduces a novel approach that integrates analytical solutions for groundwater dynamics with the U-Net (Ronneberger et al., 2015) deep learning framework to predict groundwater contaminant plume migration under dynamic pumping conditions. By incorporating the Thiem equation (Thiem, 1906) into the input preprocessing, the U-Net model transforms sparse well data into a continuous spatial field that captures the hydraulic impacts of pumping activities. This integration enables the model to leverage both deep learning capabilities and classical physics-based groundwater theories, enhancing prediction accuracy and computational efficiency. These advancements can facilitate rapid, large-scale evaluations of P&T optimization simulations, allowing for timely and effective decision-making in well placement and system management. We demonstrate the model's robust performance across both simplified transient 2D models and a more complex 3D heterogeneous site model at the 200 West P&T facility at the Hanford Site. The U-Net-based model offers substantial computational advantages, reducing simulation times significantly compared to full physics-based models and providing a powerful tool for rapid site evaluation and P&T system optimization, such as evaluating alternative P&T well network designs. Our findings highlight the potential of advanced machine learning models to significantly enhance the efficiency and sustainability of groundwater remediation efforts, offering a novel application of U-Net architecture in environmental science.

Pump-and-treat↗

Graph neural networks for CO 2 solubility predictions in Deep Eutectic Solvents

Deep Eutectic Solvents (DESs) are a promising class of solvents for CO 2 capture. DESs are complex mixtures that can be designed to optimize CO solubility and overall capture process efficiency. However, the vast design landscape of DES mixtures makes experimental investigation prohibitive; as such, there is a need for computational models that can quickly and efficiently navigate the design space and inform data collection efforts. In this work, we propose Graph Neural Network (GNN) models for predicting CO 2 solubility for DESs; the GNN leverages a mixture graph representation that captures the molecular structure of the DES components as well as their intermolecular interactions. Here, we compare the GNN framework against alternative architectures (neural networks, graph convolution networks, and random forests) and data representations (molecular fingerprints, sigma profiles, and graphs). We show that the proposed approach offers superior predictive performance; specifically, we show that solubility can be predicted reliably directly from molecular structure (without the need of using sigma profiles as proposed in previous studies). This result is important, as obtaining sigma profiles requires expensive density functional theory computations. We also explored the ability of GNNs to predict solubility for new DES mixtures and operating conditions. We found that the model extrapolates across temperature reliably. However, we also found deficiencies in the ability of the model to predict solubility for DES mixtures, pressures, and molar ratio not included in the training sets; we show that this is due to an inherent lack of chemical diversity in datasets available in the literature. The proposed computational capabilities can thus help navigate the design space of DES and inform data collection efforts. Our models, data, and benchmarks are shared as Python code implemented in Jupyter notebooks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

cclib 2.0: An updated architecture for interoperable computational chemistry

Interoperability in computational chemistry is elusive, impeded by the independent development of software packages and idiosyncratic nature of their output files. The cclib library was introduced in 2006 as an attempt to improve this situation by providing a consistent interface to the results of various quantum chemistry programs. The shared API across programs enabled by cclib has allowed users to focus on results as opposed to output and to combine data from multiple programs or develop generic downstream tools. Initial development, however, did not anticipate the rapid progress of computational capabilities, novel methods, and new programs; nor did it foresee the growing need for customizability. Here, we recount this history and present cclib 2, focused on extensibility and modularity. We also introduce recent design pivots—the formalization of cclib’s intermediate data representation as a tree-based structure, a new combinator-based parser organization, and parsed chemical properties as extensible objects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MultiTaskDeltaNet: change detection-based image segmentation for operando ETEM with application to carbon gasification kinetics

Transforming in situ transmission electron microscopy (TEM) imaging into a tool for spatially-resolved operando characterization of solid-state reactions requires automated, high-precision semantic segmentation of dynamically evolving features. However, traditional deep learning methods for semantic segmentation often face limitations due to the scarcity of labeled data, visually ambiguous features of interest, and scenarios involving small objects. To tackle these challenges, we introduce MultiTaskDeltaNet (MTDN), a novel deep learning architecture that creatively reconceptualizes the segmentation task as a change detection problem. By implementing a unique Siamese network with a U-Net backbone and using paired images to capture feature changes, MTDN effectively leverages minimal data to produce high-quality segmentations. Furthermore, MTDN utilizes a multi-task learning strategy to exploit correlations between physical features of interest. In an evaluation using data from in situ environmental TEM (ETEM) videos of filamentous carbon gasification, MTDN demonstrated a significant advantage over conventional segmentation models, particularly in accurately delineating fine structural features. Notably, MTDN achieved a 10.22% performance improvement over conventional segmentation models in predicting small and visually ambiguous physical features. This work bridges key gaps between deep learning and practical TEM image analysis, advancing automated characterization of nanomaterials in complex experimental settings.

08 HYDROGEN↗

Digital Analytics, Causal Knowledge Acquisition and Reasoning for Technical Language Processing

Complex engineering systems such as nuclear power plants (NPPs) generate and collect large amounts of equipment reliability (ER) data elements that contain information on the status of components, assets, and systems. Some of this information is textual in form and can be found in documents such as incident reports (IRs) and work orders (WOs). Analyses of textual data in current NPPs-using natural language processing (NLP) methods-have been expanded over the last decade, and it is only recently that the true potential of such analyses has emerged. So far, applications of NLP methods have mostly been limited to classification and prediction, the goal being to identify the nature of the textual element (e.g., safety or non-safety related). Here, we target a more complex problem: automatically extracting knowledge from a textual element in order to assist system engineers in conducting system health assessments. Knowledge extraction is a very broad concept, and its definition may vary depending on the application context. Our methods are a blend of both rule-based and machine learning (ML) algorithms. For our purposes, knowledge extraction means identifying the systems or assets mentioned in a given textual element, as well as the type of event described (e.g., component failure or maintenance activity). In addition, we want to capture details such as measured quantities and the temporal/cause-effect relations between events. In this tool, we also demonstrate how textual data elements are preprocessed in order to handle typos, acronyms, and abbreviations. One main feature of these methods is that they are not based solely on data, but are in fact model-based. In other words, they also rely on MBSE models that are designed to capture-from a functional point of view-the architecture of the systems/assets under consideration. The main purpose of such models is to digitally emulate system engineers' knowledge of system and asset architecture and to identify dependencies among systems, assets, and components. Provided these models, analyses of textual and numeric ER data can be performed by first identifying the OPM model elements to which the ER data elements are referring. The relationships between ER data elements are then identified by checking for any temporal or logical dependencies.

Mandelli, Diego [Idaho National Laboratory (INL), ↗

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure↗

Atomically Precise Single-Site Catalysts via Exsolution in a Polyoxometalate–Metal–Organic-Framework Architecture

Single-site catalysts (SSCs) achieve a high catalytic performance through atomically dispersed active sites. A challenge facing the development of SSCs is aggregation of active catalytic species. Reducing the loading of these sites to very low levels is a common strategy to mitigate aggregation and sintering; however, this limits the tools that can be used to characterize the SSCs. Here we report a sintering-resistant SSC with high loading that is achieved by incorporating Anderson–Evans polyoxometalate clusters (POMs, MMo 6 O 24 , M = Rh/Pt) within NU-1000, a Zr-based metal–organic framework (MOF). The dual confinement provided by isolating the active site within the POM, then isolating the POMs within the MOF, facilitates the formation of isolated noble metal sites with low coordination numbers via exsolution from the POM during activation. The high loading (up to 3.2 wt %) that can be achieved without sintering allowed the local structure transformation in the POM cluster and the surrounding MOF to be evaluated using in situ X-ray scattering with pair distribution function (PDF) analysis. Notably, the Rh/Pt···Mo distance in the active catalyst is shorter than the M···M bond lengths in the respective bulk metals. Furthermore, models of the active cluster structure were identified based on the PDF data with complementary computation and X-ray absorption spectroscopy analysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Revealing the Crystalline Architecture of Semicrystalline Ion Exchange Membranes for the Design of Conductive and Durable Alkaline Anion Exchange Membranes

Alkaline anion exchange membrane (AAEM) fuel cells offer a cost-effective alternative to proton exchange membrane (PEM) fuel cells by eliminating the need for expensive precious metal catalysts. In both PEMs and AAEMs, semicrystalline polymers are a common choice, as the crystalline domains can act as mechanical reinforcements that limit swelling and promote mechanical durability in the material. However, spatially resolved characterization of crystalline organization in ion exchange membranes beyond ensemble-averaged X-ray scattering is underrepresented, likely in part due to ionization damage limitations in soft materials. Here, in this study, we resolve the nanometer-size crystallites in semicrystalline ion exchange membranes by applying cryogenic four-dimensional scanning transmission electron microscopy (cryo-4D-STEM) along with data-processing algorithms designed to optimize signals at a low dose to minimize radiation damage. We investigate the effects of synthesis components, including molecular weight and thermal treatment, on a model system of AAEMs in comparison to Nafion, the most commonly used and commercially successful PEM today. We find that excess water uptake in polymer membranes, a property directly associated with weak mechanical durability and with possible negative impacts on ion conductivity, can be reduced by over 30% by varying the polymer's crystalline morphology through changes in synthesis parameters such as molecular weight and thermal history. Our results indicate that this improvement is correlated with smaller crystalline domains with a more homogeneous distribution. More broadly, these results demonstrate how the crystalline architecture of polymer membranes can be tuned through their chemistry and thermal treatment in order to improve their conductivity and durability for commercial fuel cell performance.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ON-OFF neuromorphic ISING machines using Fowler-Nordheim annealers

We introduce NeuroSA, a neuromorphic architecture specifically designed to ensure asymptotic convergence to the ground state of an Ising problem using a Fowler-Nordheim quantum mechanical tunneling based threshold-annealing process. The core component of NeuroSA consists of a pair of asynchronous ON-OFF neurons, which effectively map classical simulated annealing dynamics onto a network of integrate-and-fire neurons. The threshold of each ON-OFF neuron pair is adaptively adjusted by an FN annealer and the resulting spiking dynamics replicates the optimal escape mechanism and convergence of SA, particularly at low-temperatures. To validate the effectiveness of our neuromorphic Ising machine, we systematically solved benchmark combinatorial optimization problems such as MAX-CUT and Max Independent Set. Across multiple runs, NeuroSA consistently generates distribution of solutions that are concentrated around the state-of-the-art results (within 99%) or surpass the current state-of-the-art solutions for Max Independent Set benchmarks. Furthermore, NeuroSA is able to achieve these superior distributions without any graph-specific hyperparameter tuning. For practical illustration, we present results from an implementation of NeuroSA on the SpiNNaker2 platform, highlighting the feasibility of mapping our proposed architecture onto a standard neuromorphic accelerator platform.

42 ENGINEERING↗

A High-Fidelity Molecular Model of the Cu(111) Repeating Unit

Dynamic processes at surfaces are central to heterogeneous catalysis, but their atomistic mechanism(s) can prove difficult to elucidate due to variations in material structure and the corresponding impact on reactivity. Moreover, disparities between reaction conditions and those employed for spectroscopic characterization at surfaces can inhibit detailed understanding of catalysis-relevant chemistries. Herein, we substantiate the so-called “cluster-surface” analogy by leveraging a low-valent tricopper architecture ( 1 ) as a model system for small molecule activation at Cu(111). Two reaction classes are explored: the adsorption of carbon monoxide (CO) and the dissociative adsorption of dihydrogen (H 2 ). These processes serve as an ideal testbed to compare the reactivity of a molecular cluster ( 1 ) to that of a heterogeneous surface, as both reactions have empirical data from measurements performed on crystalline Cu(111). Cluster 1 reversibly binds CO. Variable temperature NMR analysis with 13 CO reveals a favorable enthalpy but large negative entropy (−5.1 kcal × mol –1 and −22.9 cal × mol –1 × K –1 , respectively) for CO binding, affording a process that is marginally endergonic at room temperature (ΔG ads (298.15 K) = 1.7 ± 0.5 kcal × mol –1 ). Similarly, analogous to a Cu(111) surface, 1 is shown to oxidatively add (chemisorb) H 2 . Kinetic parameters were determined for this process and the activation enthalpy (8.4 ± 0.5 kcal × mol –1 ) closely mirrors that established for H 2 binding at the Cu(111) facet (6.0 to 12.4 kcal × mol –1 ). Together, these results showcase that a trinuclear cluster can reproduce the small molecule binding and activation energetics of a bulk crystalline surface, setting the stage for studying less-defined surface processes in an atomically precise molecular setting.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Compressive Response and Energy Absorption of Additively Manufactured Elastomers with Varied Simple Cubic Architectures

Additive manufacturing, and particularly the vat photopolymerization process, enables the fabrication of complex geometries at high resolution and small length scales, making it well-suited for fabricating cellular structures (e.g., foams and lattices). Among these, elastomeric cellular structures are of growing interest due to their tunable compliance and energy dissipation. However, comprehensive data on the compressive behavior of these structures remains limited, especially for investigating the structure-property effects from changing the density and distribution of material within the cellular structure. This study explores how the mechanical response of polyurethane-based simple cubic structures changes when varying volume fraction, unit cell length, and unit cell patterning, which have not been systematically investigated previously in additively manufactured elastomers. Increasing volume fraction from 10% to 50% yielded significant changes in compressive stress–strain performance (decreasing strain at 0.5 MPa by 41.6% and increasing energy absorption density by 3962.5%). Although changing the unit cell length between 2.5 and 7 mm in ~30 mm parts did not result in statistically different stress–strain responses, modifying the configuration of struts of different thicknesses across designs with 30% volume fraction altered the stress–strain behavior (differences of 12.5% in strain at 0.5 MPa and 109.4% for energy absorption density). Power law relationships were developed to understand the interactions between volume fraction, unit cell length, and elastic modulus, and experimental data showed strong fits (R 2 > 0.91). These findings enhance the understanding of how multiple structural design aspects influence the performance of elastomeric cellular materials, providing a foundation for informing strategic design of tailorable materials for diverse mechanical applications.

36 MATERIALS SCIENCE↗

Autonomous Flow Electrochemistry for Accelerated Catalyst Discovery

Our objective is to develop an Autonomous Chemical Experimentation (ACE) platform that accelerates discovery of new catalytic transformations and other energy-relevant chemical reactions and processes. We intentionally designed ACE to be highly modular, both with respect to its rapid deployment to different chemistries and experimental workflows as well as incorporation of a wide range of different AI algorithms. In addition to the development of the core software architecture, initial efforts were made to incorporate Large Language Models to provide human-interpretable reasoning of the optimizer’s actions, and to develop a user-friendly graphical interface for experimental researchers. ACE was demonstrated using a flow electrocatalysis platform containing an inline FTIR spectrometer for real-time analysis and quantification of the reaction outcome. Human-in-the-loop experiments were performed in which a human researcher conducted an experiment using electrode potentials suggested by ACE, then fed the spectral data back to ACE for decision making. After confirming the successful function of the optimizer, efforts were next directed to automation of the hardware and performed full autonomy tests using three reactions: catalytic oxidation of formate, catalytic oxidation of cyclohexanol, and oxidation of hydroquinone. These studies confirm that ACE can close the loop between reaction execution, analysis, and optimization. They also reveal that more improved product detection methods will be essential for ACE to make well-informed decisions for reactions with low conversions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

Data for Promoter Deletion in the Soybean Compact Mutant Leads to Overexpression of a Gene with Homology to the C20-Gibberellin 2-Oxidase Family

Height is a critical component of plant architecture, significantly affecting crop yield. The genetic basis of this trait in soybean remains unclear. In this study, we report the characterization of the Compact mutant of soybean, which has short internodes. The candidate gene was mapped to chromosome 17, and the interval containing the causative mutation was further delineated using biparental mapping. Whole-genome sequencing of the mutant revealed an 8.7 kb deletion in the promoter of the Glyma.17g145200 gene, which encodes a member of the class III gibberellin (GA) 2-oxidases. The mutation has a dominant effect, likely via increased expression of the GA 2-oxidase transcript observed in green tissue, as a result of the deletion in the promoter of Glyma.17g145200. We further demonstrate that levels of GA precursors are altered in the Compact mutant, supporting a role in GA metabolism, and that the mutant phenotype can be rescued with exogenous GA3. We also determined that overexpression of Glyma.17g145200 in Arabidopsis results in dwarfed plants. Thus, gain of promoter activity in the Compact mutant leads to a short internode phenotype in soybean through altered metabolism of gibberellin precursors. These results provide an example of how structural variation can control an important crop trait and a role for Glyma.17g145200 in soybean architecture, with potential implications for increasing crop yield.

Biomass Analytics↗