Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Leveraging ARM Data to Improve Models for Predictive Understanding of Energy and Security Challenges

Extreme weather and natural hazards can disrupt the energy sector, affecting demand, generation, transmission, distribution, consumption and operational planning at regional and national scales. These disruptions stem from a broad range of atmospheric phenomena, including winter storms, freezing rain, wet snow loading, severe convection, flooding and landslides, wildfires, prolonged heat, and drought. Many of these same phenomena can also affect national security through impacts to transportation and infrastructure. To support the U.S. Department of Energy (DOE) focus on energy resilience and national security, the Atmospheric Radiation Measurement (ARM) User Facility is uniquely positioned to contribute measurement data, analyses, and modeling frameworks that can significantly improve predictive understanding of these hazards to mitigate their effects. To explore this opportunity, ARM convened a two-part virtual workshop in November 2025. The workshop engaged interdisciplinary experts in atmospheric science, energy systems, modeling, and operations. The goal of the meeting was to engage with these interdisciplinary experts to address three questions: • What are examples of atmospheric processes that represent significant risks to energy security or national security and where are those risks greatest? • What measurements or measurement strategies would improve ARM’s capacity to address these issues? • How can ARM and users of the ARM facility better work with the Energy Exascale Earth System Model (E3SM) and multi-sector modeling communities to apply ARM data to improving E3SM simulations of these phenomena? Participants were asked to submit white papers ahead of the meeting to initiate thinking on these themes and to help organize discussions. Workshop sessions were then organized around themes identified in the white papers. First from the white papers and then through subsequent discussions, workshop participants identified many examples that address the three questions listed above. Participants called out energy system vulnerabilities to weather phenomena such as the impact of freezing rain, strong winds, and excessive heat on power grids. They also noted the effects that weather phenomena could have on energy demand or supply (e.g., through effects of extreme temperatures). They called out security vulnerabilities such as impacts to crops from aerosol-borne pathogens and risks to industry due to melting permafrost in the Arctic. In all, over a dozen meteorological phenomena were linked to energy or security vulnerabilities. For many of the identified phenomena, participants pointed out where ARM was well poised to address issues (e.g., through measurements of cloud microphysics to inform studies of freezing rain) but also noted needs for additional measurements or modified measurement strategies. For example, adaptive scanning of severe weather would be valuable for probing winter storms or severe convection. Participants pointed out the value in integrating external observations with ARM measurements and with applying artificial intelligence (AI) to ARM observation analysis and they advocated for using model simulations to help optimize measurement strategies through Observing System Simulation Experiments (OSSEs). It was clear from the workshop that there are many ways that ARM observations can be used to mitigate energy and security concerns, but meeting participants were also asked to identify what they considered to be the greatest opportunities by ranking issues pertaining to the three workshop questions. This was accomplished through a survey administered to participants between the two virtual sessions. The highest-priority phenomena identified were winter storms, severe convection, and arctic processes. Discussion in the second session, therefore, focused primarily on these three areas, which were most fully developed in exploring ARM opportunities. Nevertheless, it was also clear that ARM has opportunities to contribute to all the identified topics. This report describes the workshop, including input from discussion and white papers (Sections 2 and 3) and a list of priority recommendations (section 4). Many other ideas for ARM contributions are discussed in individual white papers (Appendix D).

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Performance Year 1 Technical Report - OPEN COG Grid: Extendable Coherent Models-Datasets for Cognitive Power Grids

The OPEN COG Grid project is a collaborative effort between LLNL, NREL, and Texas A&M University (TAMU) to develop synthetic power system datasets that (i) contain all technical information that would be available in a real system, allowing to conduct studies ranging from dynamic simulation to long term planning studies; ii) are accessible to researchers from the broader data sciences community, as oppossed to power system experts only; and (iii) This report summarizes the work conducted during the first 15 months of execution of the project. These activities encompassed: 1. Conduct a survey of existing open data sets and open source power systems simulators, their supported use cases, and accessibility (Chapter 1). 2. Define a new extensible specification for power system data, covering all parameters necessary for most computational use cases (Chapter 2). 3. Collecting real technical system data to complete missing parameters in existing open source datasets (Chapter 3). 4. Develop models that capture the behavior of emergent actors in power grids, neglected by existing datasets; aggregated residential demand response (Chapter 4) and demand response of cryptocurrency miners (Chapter 5). 5. Collect detailed spatial information on distributed energy resources, particular, solar photovoltaic facilities (Chapter 6). The following chapters provide detailed descriptions of these tasks, the assumptions taken, and their findings. In conducting these tasks, the project team produced: two (accepted) conference papers; one journal paper under submission; one draft journal paper pending submission; released one repository with the developed power system data specification, with documentation and examples; and one extended dataset for the Texas power grid under review for release. The team hopes these contributions will enhance access to power system data and remove barriers to the development of new computational techniques for power systems, particularly, those inspired by cognitive sciences.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Comparing plant litter molecular diversity assessed from proximate analysis and 13 C NMR spectroscopy

Accurate representation of the chemical diversity of litter in ecosystem-scale models is critical for improving predictions of decomposition rates and stabilization of plant material into soil organic matter. In this contribution, we conducted a systematic review to evaluate how conventional characterization of plant litter quality using proximate analysis compares with molecular-scale characterization using 13 C NMR spectroscopy. Using a molecular mixing model, we converted chemical shift regions from NMR into fractions of carbon (C) in five organic compound classes that are major constituents of plant material: carbohydrates, proteins, lignins, lipids, and carbonylic compounds. We found positive correlations between the acid soluble fraction and carbohydrates, and between the acid insoluble fraction and lignins. However, the acid-soluble fraction underestimated carbohydrates, and the acid insoluble fraction overestimated lignins by 243%. We identified two sources of uncertainties: i) disparities between litter chemical composition based on hydrolysability and actual chemical composition obtained from NMR and ii) conversion factors to translate proximate fractions into organic constituents. Both uncertainties are critical, potentially leading to misinterpretations of decay rates in litter decomposition models. Consequently, we recommend including explicit substrate chemistry data in the next generation of litter decomposition models.

59 BASIC BIOLOGICAL SCIENCES↗

Data-driven discovery of dynamics from time-resolved coherent scattering

Coherent X-ray scattering (CXS) techniques are capable of interrogating dynamics of nano- to mesoscale materials systems at time scales spanning several orders of magnitude. However, obtaining accurate theoretical descriptions of complex dynamics is often limited by one or more factors—the ability to visualize dynamics in real space, computational cost of high-fidelity simulations, and effectiveness of approximate or phenomenological models. In this work, we develop a data-driven framework to uncover mechanistic models of dynamics directly from time-resolved CXS measurements without solving the phase reconstruction problem for the entire time series of diffraction patterns. Our approach uses neural differential equations to parameterize unknown real-space dynamics and implements a computational scattering forward model to relate real-space predictions to reciprocal-space observations. This method is shown to recover the dynamics of several computational model systems under various simulated conditions of measurement resolution and noise. Moreover, the trained model enables estimation of long-term dynamics well beyond the maximum observation time, which can be used to inform and refine experimental parameters in practice. Finally, we demonstrate an experimental proof-of-concept by applying our framework to recover the probe trajectory from a ptychographic scan. Our proposed framework bridges the wide existing gap between approximate models and complex data.

36 MATERIALS SCIENCE↗

Domain-Adaptive Neural Posterior Estimation for Strong Gravitational Lens Analysis

Modeling strong gravitational lenses is prohibitively expensive for modern and next-generation cosmic survey data. Neural posterior estimation (NPE), a simulation-based inference (SBI) approach, has been studied as an avenue for efficient analysis of strong lensing data. However, NPE has not been demonstrated to perform well on out-of-domain target data -- e.g., when trained on simulated data and then applied to real, observational data. In this work, we perform the first study of the efficacy of NPE in combination with unsupervised domain adaptation (UDA). The source domain is noiseless, and the target domain has noise mimicking modern cosmology surveys. We find that combining UDA and NPE improves the accuracy of the inference by 1-2 orders of magnitude and significantly improves the posterior coverage over an NPE model without UDA. We anticipate that this combination of approaches will help enable future applications of NPE models to real observational data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Simulations of nozzle gas flow and gas-puff Z-pinch implosions on the Weizmann Z-pinch

We present simulations of an oxygen gas puff Z-pinch on a University scale generator at the Weizmann Institute of Science. The work accounts for the detailed geometry of the nozzle, the initial neutral gas density distribution, and the subsequent implosion. The modeling results show significant improvement with data for the current at the time of stagnation in comparison with a previous effort [Rosenzweig et al., Phys. Plasmas 27, 022705 (2020)]. As a first step, we performed simulations of the flow of neutral diatomic oxygen from a plenum through a nozzle within a recessed cathode, across a gap, and into the anode with a recessed grounded honeycomb. These simulations show an agreement with the measured initial gas density profiles within the region not blocked by the recesses and accessible to visible measurements. The computed neutral gas flow profile serves as the initial condition for a radiation magnetohydrodynamic simulation of the implosion using the MACH2-TCRE code. By considering the specific details of the nozzle and chamber geometry, we find agreement with the measured current profile, including the inductive notch. The simulations predict that the plasma undergoes a strong pinch within the hidden anode recess. The simulations also predict the strongest radiation pulse occurs within the anode recess and at the time of the observed inductive notch.

Physics↗

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty guided online ensemble for non-stationary data streams in fusion science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.

AI↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

Measurements of Gamow-Teller transitions from 59 Co via the 59 Co ⁢(𝑡, 3 He +𝛾) charge-exchange reaction and its application to the stellar electron-capture rates

Electron-capture reactions on iron-group nuclei play a crucial role in the late stages of massive star evolution. Since stellar evolution simulations depend on accurate electron-capture rates—which are highly sensitive to the detailed Gamow-Teller (GT) strength distributions—reliable theoretical models are essential. However, experimental data on GT strength distributions are scarce. High-resolution measurements are therefore vital for benchmarking and improving these theoretical calculations. To provide high-resolution data on Gamow-Teller strength distributions of iron-group nuclei and to compare these results with theoretical calculations within this mass region. Differential cross sections for the 59 Co ⁢(𝑡, 3 He)⁢ 59 Fe charge-exchange reaction at 115 MeV/u were measured using the S800 spectrometer. Furthermore, to resolve individual levels that are not distinguishable in the S800 particle singles data, coincident 𝛾 rays from the 59 Fe residual nucleus were detected by using the Gamma-Ray Energy Tracking In-beam Nuclear Array 𝛾-ray tracking array. Here, the Gamow-Teller transition strength distribution from the ground state of 59 Co to 59 Fe was extracted up to an excitation energy of 10 MeV. Additionally, transition strengths for several low-lying states were determined from coincident 𝛾-ray measurements. Electron-capture rates calculated using the present data indicate that these low-lying states contribute significantly to the overall rates in relevant stellar environments. The experimental results show reasonable agreement with theoretical predictions based on both shell-model and projected shell-model calculations. High-resolution data on Gamow-Teller strength distributions—particularly for individual low-lying states—are essential for accurately determining electron-capture rates in iron-group nuclei. Coincident 𝛾-ray measurements provide a powerful tool for obtaining such detailed information. While the present work demonstrates that shell-model calculations successfully reproduce the experimental results, such comparisons are scarce and more experimental data are desirable.

59 ≤ A ≤ 89↗

Semi–Analytical Modeling of Transient Stream Drawdown and Depletion in Response to Aquifer Pumping

Analytical and semi–analytical models for stream depletion with transient stream stage drawdown induced by groundwater pumping are developed to address a deficiency in existing models, namely, the use of a fixed stream stage condition at the stream–aquifer interface. Here field data are presented to demonstrate that stream stage drawdown does indeed occur in response to groundwater pumping near aquifer–connected streams. A model that predicts stream depletion with transient stream drawdown is developed based on stream channel mass conservation and finite stream channel storage. The resulting models are shown to reduce to existing fixed–stage models in the limit as stream channel storage becomes infinitely large, and to the confined aquifer flow with a no–flow boundary at the streambed in the limit as stream storage becomes vanishingly small. The model is applied to field measurements of aquifer and stream drawdown, giving estimates of aquifer hydraulic parameters, streambed conductance, and a measure of stream channel storage. The results of the modeling and data analysis presented herein have implications for sustainable groundwater management.

54 ENVIRONMENTAL SCIENCES↗

Describing Point Defect Topology in 2D Energy Materials Through Computer Vision

Point defects such as vacancies and impurity atoms strongly impact the performance of 2D materials. Traditional efforts often rely on manual detection, a process that is time-intensive, prone to human error, and challenging to scale. Here we leverage machine learning (ML) methods to identify and quantify vacancies within 2D transition metal carbides (Ti3C2, MXenes), aiming to expedite detection while improving accuracy. MXenes exhibit valuable defect-defined electrochemical properties, but we currently lack statistical understanding of defect topology needed to fully harness these materials. Here we employ a convolutional neural network for semantic segmentation of experimental MXene images, opening an opportunity to conduct a rigorous statistical study on defect hierarchy while investigating local relaxation in the lattice. We show how the integration of ML can yield fundamental insight into point defects, providing a powerful tool that will play an increasingly crucial role in the future of materials science. ML is often not just a matter of straightforward application, and pretrained models proved ineffective in this case. Instead, we trained our own neural network (NN) and applied data augmentation techniques and fine-tuning to the training dataset. Since labeled microscopy data is often scarce, we developed training data from a previously published wide-frame MXene image, using customized Gaussian fitting to locate atomic positions. Our trained model was then applied to a large dataset of experimental images, enabling a statistical study of defect configurations across three samples prepared with different HF etchant concentrations (5%, 9.1%, and 12.5%), as shown in Fig. 1. This also allowed us to investigate local strain around vacancies, though we find that we are limited by the precision of measurements using high-angle annular dark field (HAADF) images, as shown in Fig. 2. This study demonstrates how ML enables large-scale, quantitative analysis of atomic defects - an otherwise infeasible task with traditional methods. While our NN was specialized for Ti3C2 MXenes, the pipeline we developed provides a foundation for future ML models tailored to other materials. Ultimately, we envision embedding the NN onto the microscope to give real-time feedback to the user. To make this a reality, continued work is necessary to fully understand the NN's capabilities and limitations. This study gets one step closer to our goals of automated experimentation moving away from traditional methods of manual labeling. As ML capabilities advance, we hope to continue adapting and applying these techniques in microscopy.

2D materials↗

AI Applications to Physics Experiments at Jefferson Lab

We survey how AI/ML is being deployed across Jefferson Lab's experimental and accelerator programs. In EPSCI, Hydra applies computer vision to automate real-time data-quality monitoring across all four experimental halls, replacing manual inspection of hundreds to thousands of histograms per shift. AIEC (AI Experiment Controls) uses ML to stabilize drift chamber gains and is now part of standard CEBAF production running, while AI Optimized Polarization (AIOP) targets autonomous control of polarized targets and photon beam angular alignment. In CASA, cavity fault classification models identify faulted cavities and trip types from waveform data with ~85% and ~78% agreement to labeled data, respectively, and are deployed in production; a separate effort applies LLMs and hybrid search to make the CEBAF operations logbook AI-ready. QCD-focused work includes transformer- and GAN-based generative models for particle-level event simulation, with distributed GAN training scaling studies on Polaris. Additional efforts span ML-on-FPGA for the EIC and a new Data Science Department coordinating anomaly detection, uncertainty quantification, and HPC-scalable ML lab-wide. Collectively, these projects illustrate AI's growing role in improving efficiency across JLab's nuclear physics mission.

Mei, Xinxin [Thomas Jefferson National Accelerator↗

Uncertainty-Driven Rapid Thermodynamic Assessment of Nb-Ta-Zr System and Effects of C impurities (L25GF9298S): Annual Progress Report

An integrated computational materials engineering (ICME) method is in development for rapid thermodynamic experimental investigation and high-fidelity computational modeling of refractory multi-principal element alloys (RMPEAs). These ultra-high-temperature (UHT) alloys are of interest for structural applications in extreme environments, but deficiency of reliable data, especially melting temperatures, impedes the prediction of alloys with favorable properties. The method leverages UHT capabilities and computational expertise of LLNL’s Materials Science Division and the McCormack Lab’s UHT conical nozzle levitation (CNL) system to iteratively map the Nb-Ta-Zr phase space, with focus on the liquidus surface, through targeted experiments selected by quantifying uncertainty in the thermodynamic model fitting parameters. This method will reduce the time to map uncharted RMPEA phase space and thereby accelerate discovery and development of advanced materials for applications in extreme environments.

36 MATERIALS SCIENCE↗

A New Coupled Biogeochemical Modeling Approach Provides Accurate Predictions of Methane and Carbon Dioxide Fluxes Across Diverse Tidal Wetlands

Abstract Tidal wetlands provide valuable ecosystem services, including storing large amounts of carbon. However, the net exchanges of carbon dioxide (CO 2 ) and methane (CH 4 ) in tidal wetlands are highly uncertain. While several biogeochemical models can operate in tidal wetlands, they have yet to be parameterized and validated against high‐frequency, ecosystem‐scale CO 2 and CH 4 flux measurements across diverse sites. We paired the Cohort Marsh Equilibrium Model (CMEM) with a version of the PEPRMT model called PEPRMT‐Tidal, which considers the effects of water table height, sulfate, and nitrate availability on CO 2 and CH 4 emissions. Using a model‐data fusion approach, we parameterized the model with three sites and validated it with two independent sites, with representation from the three marine coasts of North America. Gross primary productivity (GPP) and ecosystem respiration (R eco ) modules explained, on average, 73% of the variation in CO 2 exchange with low model error (normalized root mean square error (nRMSE) <1). The CH 4 module also explained the majority of variance in CH 4 emissions in validation sites ( R 2 = 0.54; nRMSE = 1.15). The PEPRMT‐Tidal‐CMEM model coupling is a key advance toward constraining estimates of greenhouse gas emissions across diverse North American tidal wetlands. Further analyses of model error and case studies during changing salinity conditions guide future modeling efforts regarding four main processes: (a) the influence of salinity and nitrate on GPP, (b) the influence of laterally transported dissolved inorganic C on R eco , (c) heterogeneous sulfate availability and methylotrophic methanogenesis impacts on surface CH 4 emissions, and (d) CH 4 responses to non‐periodic changes in salinity.

54 ENVIRONMENTAL SCIENCES↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

SITCOMTN-161: PSF assessment in the field of Abell 360 and shapeHSM shear profile using LSSTComCam data

The Rubin LSSTComCam on-sky campaign performed at the end of 2024 provided observations of the Abell 360 galaxy cluster; these data allow a preliminary study of cluster weak lensing analysis using Rubin Data Preview 1 (DP1) data. Among all the steps required for such analyses, accurate modeling of the PSF is essential. This work uses several diagnostics, mostly based on the residuals between the second moments of stars and the PSF model, to characterize the accuracy of the PSF modeling in the A360 field. We find the level of the residuals to be sufficiently low not to hinder the measurement of the tangential shear profile around A360. With a simple source selection process, we demonstrate that outputs of the LSST Science Pipelines can be used to detect the tangential shear profile in Abell 360 at the 3.6σ level, and our analysis indicates that contamination from PSF modeling systematics is negligible.

Dell'Antonio, Ian [Brown University]↗

Aboveground Biomass Estimation Using NISAR Simulated ALOS-2 Time Series Data

Aboveground biomass (AGB) is a critical parameter to better understand the global carbon cycle and to develop sustainable forest management. However, a large uncertainty prevails. L-band SAR data have demonstrated strong potential to accurately retrieve AGB over low-biomass regions (<100 Mg ha-1). The upcoming NASA-ISRO Synthetic Aperture Radar mission will collect data at L- and S-band over earth’s landmass with a repeat period of 12 days, allowing us to have ample data for monitoring biomass and its dynamics. One of the key science requirements of the mission is to produce annual AGB maps at 1-ha resolution with RMS accuracy of 20 Mg/ha for 80 percentage of area over low-biomass regions in Calibration/Validation sites. The NISAR biomass algorithm will generate AGB maps based on the parameterization of semi-empirical model along with NISAR time-series dual pol data (HH and HV). To calibrate and validate the model for mission requirements, the mission will use reference estimates of AGB produced from ground inventory plots and airborne LiDAR data collected over selected sites distributed across different global ecoregions. This paper presents the initial results of the calibration/validation of the NISAR AGB retrieval algorithm over the Lenoir Landing (LENO), Alabama, USA site using NISAR simulated ALOS-2 time series data. Five multi-temporal dual-pol HH and HV NISAR Simulated ALOS 2 data collections were used as input to assess the performance of the model. The model AGB retrieval results shows that the NISAR model was able to achieve RMS accuracy within 20 Mg/ha.

Ramachandran, Naveen [Jet Propulsion Laboratory, C↗