Search NASA⌕ Search

SEARCH · Search NASA

Results for “large-scale networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

SimH 2 : an integrated techno-economic modeling framework for hydrogen pipeline infrastructure and network optimization

Large-scale hydrogen (H 2 ) pipeline transport design and network optimization have seldom been reported due to the lack of a cost model accounting for the relationship between transport cost and hydrogen mass flow rate. Here, this work introduced a system-level cost model for hydrogen pipeline transport at supercritical state and integrated it with an existing CO 2 pipeline network tool, SimCCS, for hydrogen-specific pipeline design and optimization. The Intermountain West (I-West) region of the U.S., historically dependent on fossil fuel-based economies, is chosen to demonstrate the capabilities of our H 2 pipeline cost model and transport network optimization platform called SimH 2 . Two scenarios are examined: one where the pipeline is not allowed to pass through disadvantaged communities and the other where it is permitted. The results highlight that incorporating disadvantaged-community constraints lead to longer pipeline routes and increased transport costs, reflecting the trade-offs involved in equitable infrastructure development. It is demonstrated that the newly developed SimH 2 tool not only enables the efficient design of H 2 transportation pipelines but also optimizes the network by accounting for local terrain and the presence of disadvantaged areas.

08 HYDROGEN↗

From zonal to nodal capacity expansion planning: Spatial aggregation impacts on a realistic test-case

Solving power system capacity expansion planning (CEP) problems at realistic spatial resolutions is computationally challenging. Thus, a common practice is to solve CEP over zonal models with low spatial resolution rather than over full-scale nodal power networks. Due to improvements in solving large-scale stochastic mixed integer programs, these computational limitations are becoming less relevant, and the assumption that zonal models are realistic and useful approximations of nodal CEP is worth revisiting. Here, this work is the first to conduct a systematic computational study on the assumption that spatial aggregation can reasonably be used for ISO-scale CEP. By considering a realistic, large-scale test network based on the state of California with over 8000 buses, we find that well-designed small spatial aggregations can yield good approximations but that coarser zonal models may result in large distortions of investment decisions, e.g., capacity under-investment of up to 41% for the lowest resolution model considered.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Quantum entanglement distribution coexisting with high-rate, broadband classical optical communications over a real-world fiber connecting remote, synchronized nodes

Compatibility with existing classical network infrastructure offers a scalable path towards deploying large-scale quantum networks. Here, we demonstrate O-band polarization-encoded quantum entanglement distribution over an installed 24.4-km fiber while coexisting with a state-of-the-art fully loaded C-band classical communications line system and a picosecond-level precision L-band synchronization signal. The classical system carries two 800-Gbps channels while the remainder of the C-band is filled with amplified spontaneous emission, as is standard for such state-of-the-art communications systems. We examine the spontaneous Raman scattering spectrum generated from this broadband C-band light and offer insights into wavelength allocation for O-band quantum channels. Optimal wavelength selection and narrow filtering enable well-preserved Bell state fidelity when coexisting with 21.4-dBm aggregate launch power across the C-band suitable for 36-Tbps transmission. To the best of our knowledge, this is the first implementation of entanglement-based quantum communications between two remote nodes coexisting with independent classical communications traffic. We demonstrate coexistence of quantum entanglement with ultra-high power levels and record classical bandwidth, offering promise for real-world entanglement-based networking integrated within high-capacity communications infrastructure.

Talcott, Gina M. [Northwestern U.] (ORCID:00000002↗

From IMT Device Measurements to Network-Level Consequences: When Learning Suppresses Beyond-LIF Neuron Dynamics

Emerging neuromorphic devices such as insulator--metal transition (IMT) devices exhibit complex temporal dynamics, including slow internal state memory, hysteresis, and burst-like firing, which are poorly captured by conventional leaky integrate-and-fire (LIF) neurons. However, it remains unclear when such dynamics influence learning and inference at the network level, particularly under commonly used unsupervised plasticity rules. We present a controlled, full-stack co-design study spanning experimental characterization of individual IMT devices, compact neuron model development, and large-scale spiking network simulations with identical architectures and learning rules. Rather than optimizing benchmark accuracy, our goal is to diagnose when neuron-level dynamics survive learning and competition, and when they are suppressed, to inform the co-design of devices, networks, and learning rules that can exploit beyond-LIF complexity.

42 ENGINEERING↗

Low-energy 17 O(𝑛,𝛾)⁢ 18 O reaction within the microscopic potential model and its role for the weak 𝑟 process

The neutron radiative capture reaction 17 O ⁡(𝑛,𝛾) ⁢18 O plays a pivotal role in both nuclear structure studies and astrophysical nucleosynthesis, particularly in the formation of elements during hydrostatic and explosive stellar environments. We calculated the 17 O ⁡(𝑛,𝛾) ⁢18 O cross section within the Skyrme Hartree-Fock potential model and analyzed electric dipole 𝐸⁢1 transitions to both positive- and negative-parity states below the α-decay threshold in 18 O. Our cross sections are significantly different from the data available in commonly used libraries. We further investigate the impact of the new calculated cross section on weak 𝑟-process nucleosynthesis using large-scale reaction network calculations across a wide range of electron fractions and entropies. Our results show that the 17 O ⁡(𝑛,𝛾) ⁢18 O reaction rate significantly influences the production of first 𝑟-process peak elements, such as strontium, under specific astrophysical conditions. This study highlights the importance of accurate nuclear data for light isotopes in modeling heavy-element synthesis and provides updated reaction rates for future nucleosynthesis simulations.

6 ≤ A ≤ 19↗

Synthetic Atmospheric River Ensembles Generated by Deep-AR

This dataset contains 35,850 synthetic landfalling atmospheric river (AR) realizations generated by the Deep-AR two-stage deep-learning framework over the Northeast Pacific and U.S. West Coast. The archive contains 25 stochastic ensemble members for each of 1,434 held-out observed seed events. Each synthetic realization is initialized from conditions 48 hours before the corresponding observed AR landfall and is generated autoregressively at 6-hour intervals over a 144-hour period. Deep-AR combines a deterministic residual network (ResNet) that advances the large-scale atmospheric state with a Wasserstein generative adversarial network (WGAN) that produces stochastic, high-resolution fields. Each HDF5 file contains 0.25° gridded synthetic integrated vapor transport components (qu, qv), 10 m wind components (u10, v10), and 6-hour accumulated precipitation on a common 200 × 480 grid. The files also include coordinate and datetime arrays. This dataset supports AR hazard analysis, ensemble-based uncertainty characterization, precipitation-extremes research, and regional stress testing. Synthetic files follow the naming convention deepar.model.YYYYMMDD.HHMMSS.vNN.h5. YYYYMMDD.HHMMSS identifies the UTC initial-condition timestamp, which occurs 48 hours before the diagnosed observed landfall, and vNN identifies the zero-padded ensemble member, ranging from v01 through v25. Each synthetic file can be paired with its corresponding observed file by matching the initial-condition timestamp. The paired observed file follows the naming convention deepar.obs.YYYYMMDD.HHMMSS.h5 and is available in the separately registered oracle/deepar.obs dataset at https://wdh.energy.gov/ds/oracle/deepar.obs (DOI: https://doi.org/10.21947/3377671).

17 WIND ENERGY↗

Explainable machine learning reveals that local structural motifs encode the thermodynamic state across the CuZr metallic glass-forming range

Metallic glasses derive their properties from the statistics of local atomic motifs rather than from long-range order, yet a quantitative, chemistry-specific link between motif populations and the underlying glassy state has remained elusive. In this work we combine large-scale molecular dynamics, Voronoi tessellation, deep neural networks, and SHapley Additive exPlanations (SHAP) to identify which local structural motifs define the glassy state of Cu—Zr metallic glasses. A dataset of 17,180 atomistic configurations spanning ten compositions (Cu 20 Zr 80 –Cu 80 Zr 20 ) and four quench rates (10 9 –10 12 K/s) is used to train a feed-forward neural network that regresses temperature across the 50–2000 K liquid–supercooled–glass range, achieving a mean absolute error of 19.89 K and R 2 = 0.9974, confirming that the local structural state is faithfully encoded in motif-level structure. SHAP analysis then reveals that a tightly coupled near-icosahedral family of motifs (coordination numbers (CN) 11–13, including the full icosahedron 001200 and its single-atom-perturbation sibling 10930) collectively encodes the thermodynamic state of the system across the full glass-forming range. The CN = 11–13 ordered members carry negative SHAP values at high populations, tracking the most deeply-quenched configurations, while 10930 shows the reversed signature consistent with its role as a soft-spot host whose population shrinks as the icosahedral network deepens. The analysis demonstrates that explainable machine learning can isolate the minimal motif vocabulary defining the glassy state and recovers the near-icosahedral building blocks previously identified by data-driven analyses of Cu—Zr. The approach provides a general, chemistry-specific route for characterizing the structural state of disordered materials.

36 MATERIALS SCIENCE↗

E-PINNs: Epistemic Physics-Informed Neural Networks

Physics-informed neural networks (PINNs) have demonstrated promise as a framework for solving forward and inverse problems involving partial differential equations. Despite recent progress in the field, it remains challenging to quantify uncertainty in these networks. While techniques such as Bayesian PINNs (B-PINNs) provide a principled approach to capturing epistemic uncertainty through Bayesian inference, they can be computationally expensive for large-scale applications. In this work, we propose Epistemic Physics-Informed Neural Networks (E-PINNs), a framework that uses a small network, the epinet, to efficiently quantify epistemic uncertainty in PINNs. The proposed approach works as an add-on to existing, pre-trained PINNs with a small computational overhead. We demonstrate the applicability of the proposed framework in various test cases and compare the results with B-PINNs using Hamiltonian Monte Carlo (HMC) posterior estimation and dropout-equipped PINNs (Dropout-PINNs). In our experiments, E-PINNs achieve calibrated coverage with competitive sharpness at substantially lower cost. We demonstrate that when B-PINNs produce narrower bands, they under-cover in our tests. E-PINNs also show better calibration than Dropout-PINNs in these examples, indicating a favorable accuracy-efficiency trade-off.

AI for Science↗

Three-Dimensional Grid Visualization for Planning Activities: A Dubai Case Study

National Laboratory of the Rockies (NLR), in collaboration with the Dubai Electricity and Water Authority (DEWA) and Infra-X, has undertaken the Energy Visualization Analysis Project. The aim of this project is to enhance analytical and 3D visualization capabilities for distribution network planning and renewable energy integration. As modern grid continues to evolve with large-scale solar PV deployment and emerging distributed energy resources (DERs), the ability to effectively analyze, visualize, and communicate complex grid behaviors has become increasingly critical. The project focuses on developing empirical use cases based on real distribution feeder data and engineering workflows, ensuring the outcomes are directly aligned with operational environment. Through time-series power flow simulations and nodal hosting capacity analysis, the study quantifies the impacts of high PV penetration on voltage and thermal limits within representative 11 kV feeders. These analyses identify specific nodes and conditions where DER integration challenges arise. Furthermore, a Battery Energy Storage System (BESS) optimization algorithm was applied to determine the optimal size and placement of storage systems that can mitigate network constraints and enhance hosting capacity. The comparative results between base-case and BESS-augmented scenarios clearly demonstrate improvements in network stability and load management efficiency. In parallel, the NLR team developed an immersive 3D visualization framework, enabling interactive exploration of grid simulations using commodity head-mounted display (HMD) systems. This framework transforms conventional 2D simulation data into spatially intuitive visual environments - allowing engineers to analyze feeder conditions, PV hosting potential, and BESS effects in real time. This report represents the first foundational phase in establishing a visualization-driven analytical ecosystem. It provides a methodological foundation for data integration, visualization architecture, and simulation-based decision support, paving the way for large-scale adoption of immersive visualization across DEWA's Smart Grid Initiative, R&D activities, and future network resilience studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Simulation to a Newborn Supernova Remnant from a Low-mass Iron Core Star

Supernova remnant observations show a high degree of asymmetry, mixing, and inhomogeneity. These asymmetries are seeded during the early seconds of the explosion and are further enhanced and modified as the shock and ejecta move through the stellar progenitor and into the circumstellar medium. We present simulations of a 9.6 M⊙ zero-metallicity progenitor initialized after shock revival and evolved for several years when the ejecta is in the circumstellar medium. A suite of 1D and 2D simulations examines the effects of neutron-star wind and radioactive decay heating. In 1D, decay heating forms a low-density bubble that suppresses the reverse shock. While in 2D, the heating is localized to metal-rich pockets, inflating them and compressing the surrounding material into dense shells. In 3D, the neutron-star wind and decay heating modify the plume morphology, producing more large-scale structures. The extended plume morphology leads to an asymmetrical shock breakout. After breakout, the leading plumes cannot keep up with the shock front, resulting in deceleration and fragmentation by the reverse shock while retaining the large-scale asymmetry. The projected ejecta morphology and velocities are strongly viewing angle dependent. The relatively uniform metal-rich distribution does not resemble the strongly inhomogeneous ejecta structure of Cas A. The 160-isotope decay network shows that 24.4% of the radioactive heating comes from decay chains other than the canonical 56Ni chain. The low explosion energy, low 56Ni yield, and Ni/Fe ratio greater than unity suggest an observational signature similar to an electron capture supernova.

Neopane, Sudarshan [University of Tennessee (UT)]↗

Optimizing district energy systems by integrating Borehole Thermal Energy Storage Using a Mixed-Integer Linear Programming g-function framework with a Multi-Timescale Rolling Horizon method

Shallow geothermal has gained increasing attention in recent years; however, a reliable framework for its accurate incorporation into large-scale energy system optimization remains lacking. This study proposes a Mixed-Integer Linear Programming (MILP) framework combined with the g-function approach to integrate Borehole Thermal Energy Storage (BTES) technology into energy system optimization. Validation against a Modelica-based reservoir network simulation demonstrates that the proposed framework effectively captures the ground thermal response under varying energy loads and accurately estimates the borefield energy supply. To enhance scalability, a Rolling Horizon with Multi-Timescale (RH-MTS) method is further introduced, reducing computational time by 73 % for the 1-year optimization model with only minor loss of optimality. The framework is demonstrated through the case study of the UC Berkeley campus. Results indicate that BTES is a cost-effective and low-carbon solution: two borefields comprising 382 boreholes can meet 8.0 % and 6.6 % of the total campus heating and cooling demand, respectively, at an average energy rate of 0.70–0.77 USD/kWh and carbon intensity of 0.54 kg-CO2/kWh. Short-term analysis reveals a 35%–65% decline in BTES energy flow after 3–6 months of continuous heating/cooling operation, while long-term simulation shows that annual energy production of BTES can vary by up to 12.0 % after four years before stabilizing. Overall, this study develops a novel optimization framework that couples physics-based g-function method with MILP optimization framework, thereby advancing methodological development for shallow-geothermal integration and providing actionable guidance for BTES deployment in district-energy systems.

Yang, Jiahui↗

JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows

In modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise from multiple institutions and geographic locations are leveraged to accelerate scientific discovery. These workflows often require transmitting huge amounts of scientific data through wide-area networks. Although high-speed networks like ESnet and transfer services such as Globus have improved data mobility, several challenges remain. The sheer volume of data can overwhelm network bandwidth, widely used transport protocols such as TCP suffer from inefficiencies due to retransmissions triggered by packet loss, and existing fault-tolerance mechanisms like erasure coding introduce substantial overhead. In this paper, we propose Janus, a resilient and adaptable data transmission approach designed for cross-facility scientific workflows. Unlike traditional TCP-based methods, Janus leverages UDP, integrates erasure coding for fault tolerance, and combines it with error-bounded lossy compression to reduce overhead. This novel design allows users to balance data transmission time and accuracy, optimizing transfer performance based on specific scientific requirements. Additionally, Janus dynamically adjusts erasure coding parameters in response to real-time network conditions, ensuring efficient data transfers even in fluctuating environments. We develop optimization models for determining ideal configurations and implement adaptive data transfer protocols to enhance reliability. Through extensive simulations and real-network experiments, we demonstrate that Janus significantly improves transfer efficiency while maintaining data fidelity.

Esaulov, Vladislav [Georgia State University, Atla↗

Comparing Delay-, Distance-, and Cordon-Based Congestion Pricing Strategies Via Large-Scale Simulation

This study compares the impacts of delay-, distance-, and cordon-based congestion pricing strategies for Austin, Texas, using the POLARIS agent-based activity-based travel demand simulation model. This approach enables agent-level heterogeneity and realistic choice options (including destination, mode, and activity scheduling) for dynamic traffic assignment and congestion feedbacks across a major metro region, which are features lacking in past work. To ensure comparability, distance-based tolls were set to generate the same revenue as delay-based tolling of $3.5 M/day, averaging $1.17/resident/day or $0.42/vehicle-trip. Delay-based pricing delivers 44% lower network delay and 13% lower VHT compared to the no-toll baseline, levels unmatched by other pricing strategies. At the height of the AM peak, drivers pay up to $0.13/mile on average, though most links in the network remain untolled. Distance-based pricing is the most effective at reducing VMT (by 4%), but VHT reductions (of 6%) primarily stem from drivers selecting closer destinations, achieving only one-fourth the delay reduction of delay-based pricing. Across various implementations of delay- and distance-based pricing, the results suggest that spatial variations of tolls are far more important than temporal variations. Cordon tolls produce minimal impacts at the network-wide level, but offer substantial delay reductions inside the cordon. Other major findings include: 1) delay-based pricing increases trip-making during the PM peak period due to backward shifts in discretionary-activity start times by higher-income residents; and 2) tolls’ spatial impacts, including changes in network flows and tolls paid by residents, vary substantially between delay- and distance-based pricing strategies.

Agent-based modeling↗

Interactions Between Climate Policy and Technology-influenced Travel Behavior: Mitigating Induced Demand from CACC

Advances in vehicle technology have influenced the development of automated vehicle systems, where vehicles that do not require human intervention are already deployed in the roadway networks. While these advances are proved to increase roadway safety and highway capacity, more research is needed to understand the long-term and regional-level impacts on mobility, land use, energy consumption, and emissions. This study proposes a multi-model approach to analyze the effect of vehicle automation and deep decarbonization policies over a period from 2020 to 2040 in Austin, Texas. We use the Global Change Analysis Model (GCAM) to develop internally the scenarios that are then passed to the SMART Mobility modeling workflow, a large-scale simulation framework combining the POLARIS activity-based travel demand model and mesoscopic traffic simulator with the Autonomie vehicle energy consumption model and the UrbanSim land use simulator. Results suggest that the introduction of vehicles with advanced automation could increase fuel consumption when no decarbonization policies are implemented. Also, advances in vehicle technology research and development could lead to a decline in energy use in the long-term. Energy pricing and vehicle electrification incentives could help reduce the impact of vehicle automation. Finally, our analysis indicates the relevance of introducing land use processes in longterm vehicle automation studies.

land use↗

Risk-Aware Measurement Synchronization and Recovery for DSSE With Heterogeneous Data Sources

Power distribution systems are increasingly integrating heterogeneous sensors with varying data reporting rates and types, which pose challenges to achieving observability at the desired temporal resolution of distribution system state estimation (DSSE). Multisensor failures caused by extreme events exacerbate these issues, introducing substantial uncertainties into DSSE. This article proposes a novel solution to these challenges by ensuring high-resolution system observability despite heterogeneous data sources and multisensor failures. First, a deep learning architecture combining long short-term memory (LSTM) and graph convolutional network (GCN) is employed to synchronize meters with different reporting rates, aiming to achieve system observability. A random-walk-model-based approach is introduced to generate pseudo-measurements while properly characterizing their uncertainties under multisensor failures. Finally, a disaster-risk-informed observability metric (RiOM) is defined to quantify the uncertainty associated with state estimation results. The proposed framework offers deeper insights into the system observability on the fly compared with conventional analysis. The effectiveness of the framework is demonstrated on an IEEE standard test case and a large-scale real-world distribution feeder in mid-Minnesota in the U.S.

97 MATHEMATICS AND COMPUTING↗

Improving ProtoDUNE pion cross-section measurements with NuGraph Michel-electron tagging

Understanding hadron-argon interactions is essential for precise neutrino energy reconstruction and final-state interaction modeling in liquid-argon time projection chamber (LArTPC) experiments such as DUNE. In particular, pion absorption and charge-exchange processes constitute significant sources of systematic uncertainty in neutrino oscillation measurements. ProtoDUNE-SP, a large-scale LArTPC prototype operated at the CERN Neutrino Platform and exposed to charged-particle test beams in the few-GeV range, enables direct measurements of these processes. This work focuses on the measurement of differential cross sections for pion absorption and charge exchange using the 2 GeV/c pion beam data from the ProtoDUNE-SP run. A key component of this analysis is the identification of Michel electrons from $\pi \rightarrow \mu \rightarrow e$ decay chains, which helps separate different interaction topologies and improves background rejection. Michel electron identification will also assist in reliably calibrating the electromagnetic response in ProtoDUNE-SP data and for the future DUNE detectors. In this analysis, we apply NuGraph to identify Michel electrons. NuGraph is a graph neural network that models detector hits as nodes connected by spatial and temporal edges for particle and topology classification in LArTPC detectors. We first benchmark NuGraph’s Michel electron classification performance using ICEBERG data, a small-scale LArTPC prototype used for DUNE electronics and reconstruction development, and then transfer the approach to ProtoDUNE-SP. This poster presents the analysis strategy, NuGraph-based classification studies, and discusses how these developments are expected to improve the pion cross-section measurement.

Razafinime, Soamasina Herilala [Cincinnati U.] (OR↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Spatial Optimization of Multiscale Biorefinery Deployment for a Diversified Bioeconomy in the United States

Strategic biorefinery siting is critical for a diversified bioeconomy, yet industry, policy, and research often focus on either large-scale biofuel plants or smaller-scale specialty bioproduct facilities, with limited coordination across scales. We address this gap by modeling biorefinery deployment spanning a 28-fold difference in capacity. We developed an open-source, spatially explicit framework integrating techno-economic analysis with logistics and refinery cost surrogate models to evaluate multiscale miscanthus-derived biorefineries across the rainfed U.S. for the production of ethanol, succinic acid, lactic acid, potassium sorbate, and acrylic acid. Overall costs change little as feedstock density increases, while transport distances decrease by ∼30 to 67% (∼100 km) and siting flexibility improves. Specifically, a 5-fold feedstock density increase (2% to 10% of suitable land) reduces minimum selling prices by <10% (e.g., 0.27 USD·gal –1 for ethanol). This limited economic sensitivity suggests dense planting is not required for competitive deployment, particularly for smaller-scale facilities. Representing collection areas as irregular rather than circular expands the feasible space under low-density scenarios. While large-scale refineries anchor regional supply chains, smaller facilities retain spatial flexibility even when large refineries are established. These findings highlight the importance of spatial representation and multiscale coordination for robust, regionally tailored biomanufacturing networks to advance renewable carbon integration without extensive land conversion.

biorefinery siting↗