Search NASASearch

SEARCH · Search NASA

Results for “Large-scale Testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

From natural language to control signals: a conceptual framework for semantic channel finding in complex experimental infrastructure

Modern experimental platforms such as particle accelerators, fusion devices, telescopes, and industrial process control systems expose tens to hundreds of thousands of control and diagnostic channels, accumulated over decades of hardware evolution. Operators and AI systems alike depend on informal expert knowledge, inconsistent naming conventions, and scattered documentation to locate the signals required for monitoring, troubleshooting, and automated control, creating a persistent bottleneck for reliability, scalability, and emerging language-model-driven interfaces. We formalize semantic channel finding, the task of mapping natural-language intent to concrete control-system signals, as a general problem in complex experimental infrastructure, and introduce a four-paradigm conceptual framework to guide architecture selection based on facility-specific data regimes. The paradigms span (i) direct in-context lookup over small, curated channel dictionaries, (ii) constrained hierarchical navigation through structured trees, (iii) interactive agent exploration using iterative reasoning and tool-based database queries, and (iv) ontology-grounded semantic search that decouples channel meaning from facility-specific naming conventions. We demonstrate the practical feasibility of each paradigm through proof-of-concept implementations at four operational facilities spanning two orders of magnitude in scale: from compact free-electron lasers to large synchrotron light sources, operating under diverse control-system architectures ranging from clean hierarchical naming schemes to legacy environments with decades of heterogeneous conventions. Where evaluated against expert-curated operational queries, these instantiations achieve 90%–97% accuracy, validating the framework’s applicability across real-world deployment scenarios. To accelerate adoption across the broader scientific and industrial control-system community, we release open-source, plug-and-play implementations of all three interactive paradigms-direct lookup, hierarchical navigation, and middle-layer exploration-within the Osprey framework, together with tools for channel database generation, interactive testing, and minimal-configuration deployment. This work establishes semantic channel finding as a foundational capability for human-centric and agentic AI interfaces at large-scale facilities, providing both a systematic framework for architecture design and practical resources to enable adoption without building custom infrastructure from scratch.

channel finding

Commissioning of the large-scale lead tungstate scintillating calorimeter

Here, we report on the installation and initial commissioning of a large-scale lead tungstate (PbWO4) scintillating crystal calorimeter developed for high-rate photon detection and precise energy measurement. The calorimeter comprises 1596 high-granularity, high-resolution scintillating crystals optimized for electromagnetic-shower detection over a wide energy range. Scintillation light from each crystal is read out by Hamamatsu R4125 photomultiplier tubes equipped with a custom voltage divider and front-end amplifier to ensure stable gain at high rates. All calorimeter modules were fabricated and characterized using a light-emitting diode–based optical test system prior to installation to verify uniformity and photodetector performance. After installation, the electromagnetic calorimeter was fully integrated into the experiment data acquisition and energy-based trigger systems. The optical response of the modules was equalized using the light-monitoring system, cosmic-ray muons, and photons from Compton-scattering events. Commissioning results demonstrate a reliably calibrated optical response and stable detector performance during the first run. These results validate the calorimeter design and commissioning methodology for large-scale scintillator-based photonic instrumentation.

Analog to digital converters

Subscale Hardware-In-The-Loop Results for Hybrid Electric Turbofan Controls Use Cases

NASA is investigating hybrid electric turbine engine systems for commercial transport aircraft due to the potentially significant improvements hybrid electric technology offers in performance, fuel consumption, and operational and design flexibility. Recently, the technology has been tested at full scale in partnership with industry and advanced to Technology Readiness Level 4. This presentation will focus on a recent subscale hardware-in-the-loop test of an open source turbofan engine model developed by NASA. The Advanced Geared Turbofan 30,000 lbf – electrified (AGTF30-e) engine is used as a reference model to demonstrate control system design and use cases for an example mild hybrid electric system with no large-scale energy storage. This model is run in real-time in NASA’s Hybrid Propulsion Emulation Rig (HyPER) and is used to drive an emulation of the turbomachinery system using subscale electric machines. This dynamic scaled shaft emulation interacts with a subscale (<100 kW) hybrid system consisting of electric machines, motor controllers, and a programmable electronic load. Specific use cases demonstrated include the use of Turbine Electrified Energy Management to improve operation during transients, megawatt-scale power extraction from the AGTF30-e, and power transfer between engine spools. Results related to the effectiveness of hybrid systems are qualitatively compared to results from industry testing.

Hybrid

E-PINNs: Epistemic Physics-Informed Neural Networks

Physics-informed neural networks (PINNs) have demonstrated promise as a framework for solving forward and inverse problems involving partial differential equations. Despite recent progress in the field, it remains challenging to quantify uncertainty in these networks. While techniques such as Bayesian PINNs (B-PINNs) provide a principled approach to capturing epistemic uncertainty through Bayesian inference, they can be computationally expensive for large-scale applications. In this work, we propose Epistemic Physics-Informed Neural Networks (E-PINNs), a framework that uses a small network, the epinet, to efficiently quantify epistemic uncertainty in PINNs. The proposed approach works as an add-on to existing, pre-trained PINNs with a small computational overhead. We demonstrate the applicability of the proposed framework in various test cases and compare the results with B-PINNs using Hamiltonian Monte Carlo (HMC) posterior estimation and dropout-equipped PINNs (Dropout-PINNs). In our experiments, E-PINNs achieve calibrated coverage with competitive sharpness at substantially lower cost. We demonstrate that when B-PINNs produce narrower bands, they under-cover in our tests. E-PINNs also show better calibration than Dropout-PINNs in these examples, indicating a favorable accuracy-efficiency trade-off.

AI for Science

Fire-Resistant Textile Development for Exploration Vehicles With Enriched-Oxygen Cabin Environments

Material selection is a key part of the National Aeronautics and Space Administration (NASA) spacecraft fire safety management plan. Non-flammable textiles are necessary to ensure large-scale flame propagation events do not occur inside a spacecraft. An increased use of textiles and other softgood material is crucial to the pursuit of exploration spaceflight to reduce mass and volume. Exploration spaceflight missions benefit from enriched oxygen cabin environments (>21% O2) by allowing a reduced prebreathe protocol before Extravehicular Activity (EVA). However, materials become exponentially more flammable the higher the O2 levels become. Recent testing within the agency has revealed the lack of commercial-off-the-shelf (COTS) materials that can meet safety requirements in oxygen-enriched environments. Most of the fibers and textiles developed during Apollo (100% O2 cabin environment) and Skylab (>70% O2 cabin environment) are no longer commercially available due to those Programs ending and the discontinuation of raw materials or closure of the original manufacturers. As NASA returns to higher oxygen concentrations inside spacecraft, non-flammable textile development efforts have begun to meet the agency’s needs. This paper discusses those efforts including overall fire safety approach, priorities, interactions with industry, flammability testing and expected challenges.

Karim Aly

Fire-Resistant Textile Development for Exploration Vehicles With Enriched-Oxygen Cabin Environments

Material selection is a key part of the National Aeronautics and Space Administration (NASA) spacecraft fire safety management plan. Non-flammable textiles are necessary to ensure large-scale flame propagation events do not occur inside a spacecraft. An increased use of textiles and other softgood material is crucial to the pursuit of exploration spaceflight to reduce mass and volume. Exploration spaceflight missions benefit from enriched oxygen cabin environments (>21% O2) by allowing a reduced prebreathe protocol before Extravehicular Activity (EVA). However, materials become exponentially more flammable the higher the O2 levels become. Recent testing within the agency has revealed the lack of commercial-off-the-shelf (COTS) materials that can meet safety requirements in oxygen-enriched environments. Most of the fibers and textiles developed during Apollo (100% O2 cabin environment) and Skylab (>70% O2 cabin environment) are no longer commercially available due to those Programs ending and the discontinuation of raw materials or closure of the original manufacturers. As NASA returns to higher oxygen concentrations inside spacecraft, non-flammable textile development efforts have begun to meet the agency’s needs. This paper discusses those efforts including overall fire safety approach, state-of-the-art textile review and testing, an agency-wide assessment of textile needs performed by the NASA Engineering and Safety Center (NESC), the textile development strategy and expected challenges.

Mary Walker

Electrode Erosion and Prefire Studies Towards Fusion Scale Pulsed Power

This study presents a comprehensive investigation of electrode erosion and discharge behavior in spark gap switches over long switching cycle lifetimes. Brass, copper–tungsten (CuW), and stainless steel electrodes are tested under controlled conditions to quantify material degradation, debris accumulation, and changes in breakdown voltage. High-resolution imaging and statistical analysis of spark channel locations and gap breakdown voltages reveal how surface evolution influences long-term performance and reliability. These results provide essential data for lifetime modeling and inform design strategies for pulsed power systems in emerging applications such as private sector fusion energy and large-scale facilities like Sandia’s Z Machine and proposed ZX upgrades, where high repetition reliability and predictable behavior are critical.

electrical breakdown

Highly Efficient Selection of High-redshift Emission-line Galaxies for Future DESI-like Surveys with Deep Multiband Imaging

Emission-line galaxies (ELGs) are an important tracer of baryon acoustic oscillations (BAOs) and large-scale structure at z > 1. In this work, we investigate the feasibility of using deep wide-area multiband imaging (e.g., from the Rubin Observatory) to efficiently select high-redshift ELGs. Using Hyper Suprime-Cam grizy photometry and COSMOS2020 many-band photometric redshifts, we design simple color cuts guided by a probabilistic random forest classifier to select galaxies at z = 1.1–1.6. We then empirically test and refine these color cuts using two samples of galaxies with deep spectroscopy and broad color coverage obtained with the Dark Energy Spectroscopic Instrument (DESI). Compared to DESI ELGs at z = 1.1–1.6, we achieve a higher redshift-measurement success rate (89% versus 69%), a much higher correct redshift-range success rate (84% versus 34%), and a far higher net surface density yield (1372 deg −2 versus 660 deg −2 ). Combining our sample with current DESI ELGs would increase the net ELG number density by a factor of ∼2.5, moving it out of the shot-noise limited regime and reducing the uncertainties on the BAO scale parameter at z = 1.1–1.6 by a factor of ∼2 at the highest redshifts. We also test selections using shallower photometry and obtain qualitatively similar results.

Salcedo Hernandez, Yoquelbin [University of Pittsb

Synthetic Atmospheric River Ensembles Generated by Deep-AR

This dataset contains 35,850 synthetic landfalling atmospheric river (AR) realizations generated by the Deep-AR two-stage deep-learning framework over the Northeast Pacific and U.S. West Coast. The archive contains 25 stochastic ensemble members for each of 1,434 held-out observed seed events. Each synthetic realization is initialized from conditions 48 hours before the corresponding observed AR landfall and is generated autoregressively at 6-hour intervals over a 144-hour period. Deep-AR combines a deterministic residual network (ResNet) that advances the large-scale atmospheric state with a Wasserstein generative adversarial network (WGAN) that produces stochastic, high-resolution fields. Each HDF5 file contains 0.25° gridded synthetic integrated vapor transport components (qu, qv), 10 m wind components (u10, v10), and 6-hour accumulated precipitation on a common 200 × 480 grid. The files also include coordinate and datetime arrays. This dataset supports AR hazard analysis, ensemble-based uncertainty characterization, precipitation-extremes research, and regional stress testing. Synthetic files follow the naming convention deepar.model.YYYYMMDD.HHMMSS.vNN.h5. YYYYMMDD.HHMMSS identifies the UTC initial-condition timestamp, which occurs 48 hours before the diagnosed observed landfall, and vNN identifies the zero-padded ensemble member, ranging from v01 through v25. Each synthetic file can be paired with its corresponding observed file by matching the initial-condition timestamp. The paired observed file follows the naming convention deepar.obs.YYYYMMDD.HHMMSS.h5 and is available in the separately registered oracle/deepar.obs dataset at https://wdh.energy.gov/ds/oracle/deepar.obs (DOI: https://doi.org/10.21947/3377671).

17 WIND ENERGY

Stacked reverberation mapping of high-redshift quasars in DESI. I. Feasibility analysis

The broad-line region of quasars has long been probed by reverberation mapping techniques that measure time lags between continuum and broad emission-line variations. Stacked reverberation mapping has been proposed as a less observationally expensive alternative to traditional methods. This ensemble approach also reduces biases from small-number statistics. The Dark Energy Spectroscopic Instrument (DESI) is conducting the most extensive spectroscopic survey of quasars to date. We create mock light curves emulating expected DESI quasar observations at redshifts $1.48\lt z\lt 5.2$ and luminosities $44.68 \le \log \lambda L_{1350 \mathring{\rm A}{}} / \mathrm{erg\, s^{-1}} \le 45.99$ to test stacked reverberation mapping feasibility using sparse spectroscopic data paired with well-sampled photometric data. The pipeline, using the lag estimation code JAVELIN (Just Another Vehicle for Estimating Lags In Nuclei), successfully recovers the simulated C IV lags within 1σ of the true values using spectroscopic light curves composed of only a few spectral epochs (2–10) with irregular cadences. We investigate how observational factors, including C IV flux error magnitude, number of stacked quasars, and spectral epoch count, affect performance. This work motivates a pathway for future stacked reverberation mapping projects with large-scale spectroscopic surveys of quasars having $\ge 2$ spectroscopic observations. Our results suggest an economical alternative for constraining and extending the radius–luminosity relation to higher redshifts and luminosities. Subsequently, this relation can be employed more reliably in single-epoch black hole mass measurements and quasar cosmology in these distant regimes.

quasars: general, quasars: supermassive black hole

An agentic artificially intelligent X-ray scientist

Executing experimental tasks in both normal research laboratories and large-scale scientific facilities often requires extensive human supervision and remains a key challenge on the path to fully autonomous, artificial intelligence (AI)-driven science. Here we demonstrate a large language model-driven agent that autonomously performs X-ray sample alignment on a synchrotron beamline by planning actions, executing instrumental commands, interpreting observations and iterating towards experimental goals. Based on existing large language models with structured tool-use via the model context protocol, our AI X-ray scientist was guided and tested using an in-house-built virtual experimental setup that mirrors a six-circle diffractometer at an operational synchrotron beamline. The agentic workflow developed in the virtual environment was directly deployed on a real beamline, where it correctly identified reference reflections and determined the orientation matrix, an essential first step in any type of single-crystal scattering experiment. Our AI X-ray scientist responded effectively to unexpected experimental conditions, demonstrating adaptive problem-solving and readiness for addressing practical experimental situations. Our study provides a step towards autonomous operation across diverse experimental environments at large-scale scattering facilities.

Chen, Zhantao (ORCID:0000000319543868)

Ensemble Kalman filter for data assimilation coupled with low-resolution computations techniques applied in fluid dynamics

This paper presents an innovative Reduced-order model (ROM) for merging experimental and simulation data using data assimilation (DA) to estimate the "True" state of a fluid dynamics system, leading to more accurate predictions. Our methodology introduces a novel approach by implementing the ensemble Kalman filter (EnKF) within a reduced-dimensional framework, grounded in a robust theoretical foundation and applied to fluid dynamics. To address the substantial computational demands of DA, the proposed ROM employs low-resolution (LR) techniques to drastically reduce computational costs. This innovative approach involves downsampling datasets for DA computations, followed by an advanced reconstruction technique based on low-cost singular value decomposition (lcSVD). The lcSVD method, a key innovation in this paper, has never been applied to DA before and offers a highly efficient way to enhance resolution with minimal computational resources. Our results demonstrate significant reductions in both computation time and RAM usage through these LR techniques without compromising the accuracy of the estimations. For instance, in a turbulent test case, for a data compression rate of 15.9, the LR approach can achieve a speed-up of 13.7 and a RAM compression of 90.9% while maintaining a low relative root mean square error (RRMSE) of 2.6%, compared to 0.8% in the high-resolution (HR) reference. Furthermore, we highlight the effectiveness of the EnKF in estimating and predicting the state of fluid flow systems based on limited observations and given low-fidelity numerical data. This paper highlights the potential of the proposed DA method in fluid dynamics applications, particularly for improving computational efficiency in CFD and related fields. Its ability to balance accuracy with low computational and memory costs makes it especially suitable for large-scale and real-time applications, such as environmental monitoring or engineering design. This method will be incorporated into ModelFLOWs-app.

Data Assimilation

A 3D-Printed Millimeter-Wave Inline Waveguide-to-Coplanar-Waveguide Transition to Enable Dense Spectrometer Arrays for Intensity Mapping Surveys

We present a 3D-printed millimeter-wave, octave-bandwidth, in-line waveguide-to-coplanar-waveguide transition designed to enable focal planes with dense arrays of on-chip spectrometers. These arrays will enable compelling surveys of the large-scale structure of the universe through millimeter-wave intensity mapping. The transition consists of a four-step ridge-waveguide transformer that couples light from a rectangular waveguide onto a coplanar waveguide via an electrical connection made with indium bump bonds. We develop a tolerance-aware optimization approach to identify high-performance transition geometries that are robust to manufacturing variations; the same formulation can be applied to other tolerance-sensitive design problems. We also describe the implementation of a custom apparatus and procedure for bump-bonding a silicon chip to a metallized 3D-printed component. We detail the fabrication of the coplanar waveguide chip and three-dimensional waveguide structure, simulations and metrology of a test device, and room temperature reflectance measurements of this device. The room temperature metrology and reflection measurements are consistent with a model that predicts a coupling efficiency of $\mathord{\sim} 95\%$ at cryogenic temperatures in the 85-170 GHz frequency range.

Stover, Austin [Chicago U.; Chicago U., KICP] (ORC

Dark Energy Survey year 6 results: Magnification modeling and its impact on galaxy clustering and galaxy-galaxy lensing cosmology

Gravitational lensing magnification alters the observed spatial distribution of galaxies and must be accounted for to prevent biases in cosmological probes of the large-scale structure. We investigate its effects on the Dark Energy Survey Year 6 galaxy clustering and galaxy-galaxy lensing analyses using the fiducial lens (position tracer) sample M ag L im++. Magnification bias is parameterized by a coefficient that describes the response of the number of selected objects per unlensed area element to a change in the lensing convergence. We quantify this coefficient using the BALROG synthetic source injection catalog to account for the complexity of the selection function, and compare these results with simplified estimates. The resulting values of the magnification coefficients for each redshift bin are [3.16 ± 0.08, 2.76 ± 0.21, 4.09 ± 0.15, 4.42 ± 0.16, 4.90 ± 0.29, 4.83 ± 0.25]. Relative to Year 3, this analysis provides more precise and accurate magnification bias estimates through a larger BALROG area and reweighting to better match the data properties. Here, the cosmological results are robust when tested against various magnification parameter prior choices and also when adding cross-clustering between lens redshift bins. Neglecting magnification, however, introduces significant systematic shifts: relative to the fiducial analysis with Gaussian priors centered on the BALROG -derived estimates, we observe shifts of 1.37σ in S 8 and -0.84σ in Ω m (with cosmic shear included: -0.61σ in S 8 and -0.71σ in Ω m ), in agreement with findings from simulated data, demonstrating that magnification must be modeled to avoid biases. Freeing the magnification bias in lens bin 2 leads to unphysical negative values, further justifying its exclusion from the fiducial Year 6 analysis.

Cosmological parameters

Robustness of pairwise kinematic Sunyaev–Zel’dovich effect to optical-cluster-selection bias

The pairwise kinematic Sunyaev–Zel’dovich (kSZ) effect measures both the pairwise motion between galaxy groups and clusters and the amount of gas within them, providing a tracer for cosmic growth. To interpret the cosmological information in the kSZ measurements, it is crucial to understand the optical-cluster-selection bias on the kSZ observables. Line-of-sight structures that contribute to both the optical observable (e.g., richness) and the cosmological signal can induce a correlation between these two quantities at a fixed cluster mass. The selection bias arising from this correlation is a key systematic effect for cosmological analyses. For cosmological observables such as cluster abundance and weak lensing, controlling this selection bias may help explain the tension between the DES-Y1 results and the Planck constraints. In order to test for a kSZ effect equivalent of such a bias, we adopted an alternative mock richness based on galaxy counts within cylindrical volumes along the line of sight. We applied the cylindrical count method to hydrodynamical simulations across a wide range of galaxy-selection criteria, assigning richness consistent with DES-Y1 to the mock clusters. When comparing optically selected clusters to mass-selected halos, we find no significant bias on pairwise kSZ signals, pairwise velocities, or optical depth within our uncertainty limits of approximately 16, 10, and 8%, respectively.

79 ASTRONOMY AND ASTROPHYSICS

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE

Nodal capacity expansion planning with flexible large-scale load siting

We propose explicitly incorporating large-scale load siting into a stochastic nodal power system capacity expansion planning model that concurrently co-optimizes generation, transmission, and storage expansion. The potential operational flexibility of some of these large loads is also taken into account by considering them as consisting of a set of tranches with different reliability requirements, which are modeled as a constraint on expected served energy across operational scenarios. We implement our model as a two-stage stochastic mixed-integer optimization problem with cross-scenario expectation constraints. To overcome the challenge of scalability, we build upon existing work to implement this model on a high performance computing platform and exploit scenario parallelization using an augmented Progressive Hedging Algorithm. The algorithm is implemented using the bounding features of mpisppy, which have shown to provide satisfactory provable optimality gaps despite the absence of theoretical guarantees of convergence. We test our approach and assess the value of this proactive planning framework on total system cost and reliability metrics using realistic testcases geographically assigned to San Diego and South Carolina, with datacenter and direct air capture facilities as large loads.

24 POWER TRANSMISSION AND DISTRIBUTION

Final Technical Report for U.S.-Japan Hadronic Physics Exchange Program for Studies of Hadron Structure and QCD

Nuclear physics explores the fundamental properties of matter -- how protons and neutrons emerge as quantum systems of elementary particles, how they form the atomic nuclei, and how they give rise to the wide variety of phenomena and applications at biological, technical, and astronomical scales. It is a global scientific effort centered around large-scale experimental user facilities (particle accelerators and detectors), advanced theoretical methods and concepts, and computational techniques and resources. Exchange of knowledge and ideas, scientific collaboration, and workforce development on a global scale are essential for the future of the field. The nuclear physics program envisaged in the 2023 DOE/NSF NSAC Long-Range Plan and pursued at the U.S. National Labs has strong synergies with programs at other facilities worldwide and will realize significant benefits from international collaboration. Nuclear physics is also recognized for promoting international cooperation in the broadest sense through joint construction and operation of experimental equipment, personal contacts between scientists, and education and training. The U.S.-Japan Hadronic Physics Exchange Program (USJPHE) supported collaborative scientific research in hadronic physics and quantum chromodynamics. USJHPE focused on subject areas related to the programs at current and future experimental facilities in the U.S.\ and Japan and supported both experimental and theoretical studies. USJHPE particularly aimed to realize synergies between the hadronic physics programs at Jefferson Lab 12 GeV and J-PARC resulting from the complementarity of electromagnetic and hadronic probes in the multi-GeV energy range. Subject areas of common interest included the quark-gluon structure of hadrons and nuclei, meson and baryon spectroscopy, strangeness and hypernuclear physics, and other related topics. USJHPE also supported research in hadronic physics and nuclear-physics-enabled tests of fundamental symmetries related to the programs at Brookhaven National Lab, Fermilab, KEK, Spring-8, and university-based facilities in the U.S. and Japan. USJHPE especially promoted collaboration between the U.S. and Japanese nuclear physics communities in developing the physics program and instrumentation for the future Electron-Ion Collider. USJHPE was intended to provide travel grants to U.S.-based scientists (primary institutional affiliation with a U.S.\ university, national laboratory, or other research center) to visit Japanese institutions and conduct collaborative research there. The program supported senior researchers, postdoctoral fellows, and students. Continuing the setup of the preceding grant period, J-PARC served as the Japanese “hub” for U.S. physicists for short- and long-term visits, and JLab served as the corresponding U.S. “hub”. The program was officially managed through the U. of Connecticut in Storrs, CT. Support for Japanese physicists visiting the U.S. was provided through funds from Japanese funding agencies. The USJHPE program promoted the scientific exchange and the collaborative spirit in hadronic physics between the two countries.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS