Search NASA⌕ Search

SEARCH · Search NASA

Results for “semi-empirical model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Deep-learning atomistic semi-empirical pseudopotential model for nanomaterials

The semi-empirical pseudopotential method (SEPM) has been widely applied to provide computational insights into the electronic structure, photophysics, and charge carrier dynamics of nanoscale materials. We present “DeepPseudopot”, a machine-learned atomistic pseudopotential model that extends the SEPM framework by combining a flexible neural network representation of the local pseudopotential with parameterized non-local and spin-orbit coupling terms. Trained on bulk quasiparticle band structures and deformation potentials from GW calculations, the model captures many-body and relativistic effects with very high accuracy across diverse semiconducting materials, as illustrated for silicon and group III-V semiconductors. DeepPseudopot’s accuracy, efficiency, and transferability make it well-suited for data-driven in silico design and discovery of novel optoelectronic nanomaterials.

Lin, Kailai [University of California, Berkeley, C↗

Characterization of uncertainties in electron-argon collision cross sections

Abstract The predictive capability of a plasma discharge model depends on accurate representations of electron-impact collision cross sections, which determine the corresponding reaction rates and electron transport properties. The values of cross sections can be known only approximately either through experiments or simulations and are thus subject to uncertainties. Quantifying the uncertainties in plasma simulations allows us to assess the reliability of simulations and to provide a basis for interpreting discrepancies between simulations and experiments. For such uncertainty quantification of plasma simulations, it is essential to quantify the uncertainties of the underlying cross sections. Although much effort has been committed to calibrate the cross section values, their uncertainties are not well investigated. We characterize uncertainties in electron-argon atom collision cross sections using a Bayesian framework. Six collision processes—elastic momentum transfer, ionization, and four excitations—are characterized with semi-empirical models, which effectively capture the features important to the macroscopic properties of the plasma. A probability model for the uncertain parameters of these semi-empirical models is developed. Specifically, a Gaussian-process likelihood model is proposed to capture discrepancies among data sets, as well as the model-form inadequacies of the semi-empirical models. Two other likelihood models are compared with the proposed Gaussian-process model, to illustrate the importance of the choice of the likelihood model. The cross section models are calibrated using the electron-beam experiments and ab-inito quantum simulations. The resulting calibrated uncertainties capture well the scattering among the data sets. The calibrated cross section models are further validated against swarm-parameter experiments and zero-dimensional Boltzmann equation simulations of widely used cross section datasets.

Chung, Seung Whan (ORCID:0000000302501549)↗

Scaling dynamics in low-salt-rejection reverse osmosis for high-salinity produced water desalination: Mechanistic modeling and membrane autopsy

Membrane scaling remains a critical barrier to the reliable operation of desalination systems, particularly for hypersaline produced water (PW) treatment. This study fills the knowledge gap of autopsy-based model validation for PW desalination by elucidating scaling mechanisms in a Low-Salt-Rejection Reverse Osmosis (LSRRO) system through the integration of pilot-scale experimentation and complementary modeling approaches. A semi-empirical modeling framework was developed and applied to a multistage pilot LSRRO system equipped with nanofiltration and RO membranes treating high-salinity PW from the Permian Basin. Water quality analysis showed that total dissolved solids decreased from ~130,000 mg/L to ~1900 mg/L in the permeate, then further reduced to ~300 mg/L by a second-pass RO. Two different thermodynamic modeling approaches were evaluated: the first extends the LSRRO framework by incorporating system complexity and scaling phenomena, whereas the second method explicitly captures concentration polarization in localized supersaturation. Both methods illustrate the tendency for carbonate and sulfate scaling throughout the stages. Membrane autopsies revealed a silica-dominated deposit matrix, localized CaSO 4 at Stage 2, and minor barite/celestite despite their prominence in model predictions. Quantum-chemical calculations indicated silica scaling can be rationalized by favorable adsorption of H 4 SiO 4 on Fe-oxide surfaces (ΔG ≈ −44 kJ/mol), providing a kinetic pathway for interfacial inorganic polymerization even when bulk equilibrium predictions are conservative. Overall, the thermodynamic scaling modeling and membrane autopsy revealed heterogeneous, localized deposits with limited impact on LSRRO performance, while quantum analysis rationalized the thermodynamically unfavorable precipitation formation under bulk equilibrium, reconciling model–autopsy discrepancies. These insights support targeted pretreatment and silica-specific antiscalants to extend membrane lifetime and increase recovery, providing a transferable framework for hypersaline water desalination systems. The combined experimental–computational approach provides new mechanistic insight into scaling in hypersaline membrane systems and establishes a transferable framework for predicting and mitigating scaling in next-generation desalination technologies.

Low-salt-rejection reverse osmosis↗

Stellar Mass Calibrations for Local Low-mass Galaxies

The stellar masses of galaxies are measured from integrated light via several methods—however, few of these methods were designed for low-mass (M ⋆ ≲ 10 8 M ⊙ ) “dwarf” galaxies, whose properties (e.g., stochastic star formation, low metallicity) pose unique challenges for estimating stellar masses. In this work, we quantify the precision and accuracy at which stellar masses of low-mass galaxies can be recovered using UV/optical/IR photometry. We use mock observations of 469 low-mass galaxies from a variety of models, including both semi-empirical models (GRUMPY and UniverseMachine-SAGA) and cosmological baryonic zoom-in simulations (MARVELous Dwarfs and FIRE-2), to test literature color–M ⋆ /L relations and multiwavelength spectral energy distribution (SED) mass estimators. We identify a list of “best practices” for measuring stellar masses of low-mass galaxies from integrated photometry. We find that literature color–M ⋆ /L relations are often unable to capture the bursty star formation histories (SFHs) of low-mass galaxies, and we develop an updated prescription for stellar mass based on g − r color that is better able to recover stellar masses for the bursty low-mass galaxies in our sample (with ∼0.1 dex precision). SED fitting can also precisely recover stellar masses of low-mass galaxies, but this requires thoughtful choices about the form of the assumed SFH: Parametric SFHs can underestimate stellar mass by as much as ∼0.4 dex, while nonparametric SFHs recover true stellar masses with insignificant offset (−0.03 ± 0.11 dex). Finally, we also caution that noninformative (wide) dust attenuation priors may introduce M ⋆ uncertainties of up to ∼0.6 dex.

de los Reyes, Mithi A. C. [Amherst College, MA (Un↗

First principles free energy model with dynamic magnetism for δ -plutonium

We present an ab initio free energy model derived from a fully relativistic density functional theory (DFT) electronic structure with dynamic magnetism for δ -plutonium (face-centered cubic, fcc). The DFT model is extended with orbital-orbital interaction in a parameter free orbital polarization (OP) mechanism consistent with previous modeling of plutonium. Gibbs free energy is built from components associated with the temperature dependence of the electronic structure and the corresponding electronic entropy, lattice vibrations within an anharmonic lattice dynamics model, and dynamical fluctuations of the magnetization density, i.e. magnetic fluctuations. The fluctuation model consists of transverse and longitudinal modes driven by temperature induced excitations of the DFT + OP electronic structure. The ab initio model thus incorporates fluctuating states beyond the electronic ground state. Thanks to the dynamic magnetism, the theory predicts excellent thermodynamic properties and a Gibbs free energy in accord with CALPHAD and semi-empirical modeling developed from the thermodynamic observables. The magnetic fluctuations further explain anomalous behaviors of the thermal expansion in plutonium. Specifically, a thermal expansion for the δ -plutonium system turning from positive to negative at temperatures above room temperature, a tendency for gallium to reduce and remove the negative thermal expansion depending on composition, and a positive thermal expansion for the high temperature ϵ phase.

dynamic↗

Decoupling electrode kinetics to elucidate reaction mechanisms in alkaline water electrolysis

Alkaline water electrolysis (AWE) presents key advantages, including reduced material costs, enhanced operational stability, and compatibility with non-precious metal catalysts, positioning it as a scalable route for hydrogen production. In this study, we introduce a minimally invasive single-cell configuration incorporating a reference electrode via diaphragm extension to form an internal ion channel. This setup, combined with an interfaced potentiostat and auxiliary electrometer, enables real-time, independent monitoring of anode and cathode behavior, offering high-resolution electrochemical diagnostics. While it is well established that the hydrogen evolution reaction (HER) exhibits sluggish kinetics in alkaline media, our study reveals that this limitation persists even in practical AWE systems where nickel-based substrates are used as electrodes. This observation is supported by both experimental data and voltage breakdown modeling. Arrhenius-type analysis reveals that localized electric fields induced by catalysts shift the reaction kinetics from classical Butler–Volmer behavior toward a Marcus-like regime, where interfacial molecular dynamics and bimolecular charge transfer dominate. We propose a semi-empirical model and a surficial reaction mechanism to describe these dynamics. This work underscores the critical need for cathode innovation and provides a rational framework for designing advanced catalysts and electrode architectures to optimize AWE performance.

08 HYDROGEN↗

Many-body expansion based machine learning models for octahedral transition metal complexes

Abstract Graph-based machine learning (ML) models for material properties show great potential to accelerate virtual high-throughput screening of large chemical spaces. However, in their simplest forms, graph-based models do not include any 3D information and are unable to distinguish stereoisomers such as those arising from different orderings of ligands around a metal center in coordination complexes. In this work we present a modification to revised autocorrelation descriptors, a molecular graph featurization method, for predicting spin state dependent properties of octahedral transition metal complexes (TMCs). Inspired by analytical semi-empirical models for TMCs, the new modeling strategy is based on the many-body expansion (MBE) and allows one to tune the captured stereoisomer information by changing the truncation order of the MBE. We present the necessary modifications to include this approach in two commonly used ML methods, kernel ridge regression and feed-forward neural networks. On a test set composed of all possible isomers of binary TMCs, the best MBE models achieve mean absolute errors (MAEs) of 2.75 kcal mol −1 on spin-splitting energies and 0.26 eV on frontier orbital energy gaps, a 30%–40% reduction in error compared to models based on our previous approach. We also observe improved generalization to previously unseen ligands where the best-performing models exhibit MAEs of 4.00 kcal mol −1 (i.e. a 0.73 kcal mol −1 reduction) on the spin-splitting energies and 0.53 eV (i.e. a 0.10 eV reduction) on the frontier orbital energy gaps. Because the new approach incorporates insights from electronic structure theory, such as ligand additivity relationships, these models exhibit systematic generalization from homoleptic to heteroleptic complexes, allowing for efficient screening of TMC search spaces.

Meyer, Ralf (ORCID:0000000322360261)↗

Assessment of metadynamic recrystallization in single copper particle impacts by focused ion beam tomography

We study single Cu-on-Cu impacts relevant to cold spray deposition and quantitatively analyze the metadynamic recrystallization (mDRX) that takes place after the impact by virtue of lingering impact adiabatic heat. Unlike prior studies, the current full 3D tomographic analysis of the mDRX volume shows that mDRX is extremely common in such impacts, although it is often missed when examining 2D sections. We also report an unexpected trend: there is a “sour spot” for mDRX at velocities about 20–40 % above the velocity for particle adhesion. This non-monotonic trend is contrary to the expectations based on increasing adiabatic heating with velocity. With a schematic model, we show that the trend can be explained on the basis of heat transfer: cooling of the heat-affected region is limited by transport through the bonded regions at the particle-substrate interface. Thus, bonding has a prominent role in the heat dissipation process and the best bonded particles most rapidly bulk quench, avoiding mDRX. Here, the developed semi-empirical model aligns with the experimental findings and may help inform microstructural evolution during cold spray and post-spray processing.

FIB-SEM tomography↗

Compressibility and permeability of particulated non-recyclable municipal solid waste

Biofuels from non-recyclable municipal solid waste (NMSW) stand at the forefront of energy sustainability. However, their widespread adoption is hampered by persistent material handling issues stemming from the variability in NMSW material properties. An enhanced understanding of particulated NMSW properties, particularly compressibility and permeability, is essential to address the feedstock handling challenges and optimize the handling equipment. This study measures the compressibility and gas permeability of five streams of NMSW materials (i.e., rigid plastics, cardboard, thin film, paper, and foam) and their mixtures under different stress conditions. The results highlight the significant variability in compressibility and gas permeability among different streams, as well as the impacts of particle sizes. A semi-empirical model capable of predicting the gas permeability of NMSW mixtures is established and validated. Here, the results also highlight that reduced NMSW particle size helps promote consistency in NMSW feedstock’s physical and mechanical properties, which is favored for handling equipment design in waste-to-energy recovery facilities.

09 BIOMASS FUELS↗

Time-Domain Vortex Induced Vibration Modeling of Reference Dynamic Power Cable for the Gulf of Maine

Vortex-induced vibration (VIV) is a phenomenon known to decrease the fatigue life of dynamic power cables used in floating offshore wind systems through increased bending loading cycles. A variety of modeling solutions have been proposed to study this, but none enable fully-coupled simulations leveraging open-source tools. To address this, a time-domain VIV model has been successfully implemented into the open-source mooring dynamics model MoorDyn. The semi-empirical VIV model describes a lift force acting on segments of a flexible cylinder, making it well suited for implementation into MoorDyn. The new capability successfully predicted the frequencies and magnitudes of strain in both steady and oscillating flows when compared to the original validation results and matched peak spectral response when compared to experimental results, verifying successful implementation. MoorDyn with VIV was then leveraged to simulate a 15 MW floating turbine with a dynamic power cable in the Gulf of Maine using OpenFAST for six return periods of combined wind, wave, and current conditions. Increased curvature and tension fluctuations in higher flow speeds were observed when simulating VIV. Maximum tensions and curvatures also increased, with the 500 year conditions violating the curvature factor of safety of 2.0. These results highlight the importance of considering VIV when designing dynamic power cables for the Gulf of Maine, and indicate that cost effective mitigation strategies should be explored for the area. They also demonstrate the utility of this new modeling capability, which provides the first open-source tool for fully coupled time-domain floating offshore wind simulations with dynamic power cable VIV.

17 WIND ENERGY↗

Daily operational impacts on battery degradation in heavy-duty electric drayage trucks

Battery aging is a critical factor influencing the performance, longevity, and cost of ownership of battery electric trucks (BETs). This paper presents a comprehensive evaluation of battery aging for two Li-ion battery chemistries, Nickel-Manganese-Cobalt (NMC) and Lithium-Iron-Phosphate (LFP), accounting for both cycling and calendar aging. In contrast to traditional methods that rely on simplified linear degradation models based on manufacturer-provided data, this study employs semi-empirical aging models calibrated to experimentally collected data. The models are integrated into a detailed vehicle simulation environment, enabling a comprehensive assessment of battery degradation under realistic operating conditions. A case study focusing on heavy-duty electric drayage truck operations in the Port of Savannah, GA, is presented to illustrate the impact on battery pack lifespan of: seasonal variations, daily operational activities, charging strategies, and battery storage conditions. The results illuminate the significance of the battery pack’s state of charge during stationary periods, such as overnight storage or weekend parking, on battery degradation and its potential implications for long-term vehicle viability. Additionally, the study explores how different operational and environmental factors affect battery degradation, offering critical insights into best battery charging and storage practices. Our results demonstrate that LFP outperforms NMC in terms of years of useful life; however, by utilizing charging strategies that minimize the amount of time the battery spends resting at high levels of state-of-charge, the lifespan of the battery pack that uses NMC can nonetheless be increased by more than a factor of two.

25 ENERGY STORAGE↗

Modeling the effects of active wake mixing on wake behavior through large-scale coherent structures

The use of active wake mixing (AWM) to mitigate downstream turbine wakes has created new opportunities for reducing power losses in wind farms. However, many current analytical or semi-empirical wake models do not capture the flow instabilities that are excited through the blade pitch actuation. In this work, we develop a framework, which accounts for the impacts of the large-scale coherent structures and turbulence on the mean flow, for modeling AWM. The framework uses a triple-decomposition approach for the unsteady flow field and models the mean flow and fine-scale turbulence with a parabolized Reynolds-averaged Navier–Stokes (RANS) system. The wave components are modeled using a simplified spatial linear stability formulation that captures the growth and evolution of the coherent structures. Comparisons with high-fidelity large eddy simulations (LESs) of the turbine wakes showed that this framework was able to capture the additional wake mixing and faster wake recovery in the far-wake regions for both the pulse and helix AWM strategies with minimal computational expense. In the near-wake region, some differences are observed in both the RANS velocity profiles and initial growth of the large-scale structures, which may be due to some simplifying assumptions used in the model.

17 WIND ENERGY↗

Modelling Tritium Production and Release at High-Energy Proton Accelerators

Tritium is a well-known byproduct of particle accelerator operations. To keep levels of tritium below regulatory limits, tritium production is actively monitored and managed at Fermilab. We plan to study tritium production in the targets, beamline components, and shielding elements of the Fermilab facilities such as NuMI, BNB, and MI-65. To facilitate the analysis, we construct a simple model and use three Monte Carlo radiation codes, FLUKA, MARS, and PHITS, to estimate the amount of tritium produced in these facilities. The analysis could also serve as an intercomparison between these code results related to tritium production. To assess the actual amounts of tritium that would be released from various materials, we employ a semi-empirical diffusion model. The results of this analysis are compared to experimental data whenever possible. This approach also helps to optimize proposed target materials with respect to the tritium production and release.

Georgobiani, Dali [Fermilab]↗

Next Generation Heat Transfer Fluids for Two-Phase Immersion Cooling of Data Centers

The purpose of this study is to evaluate the performance of next generation dielectric fluids in a Two-Phase Immersion Cooling (2PIC) system, which was designed for use in data centers. Hence, this report contains the performance evaluations of a new developmental dielectric fluid, Opteon™ 2P50, in a commercially available small-scale 2PIC system under typical and off-design range of operating conditions. Accordingly, ambient temperature and thermal loads were varied to simulate different ambient conditions. Additionally, this research report describes the development of a semi-empirical lumped model to predict the energy efficiency of the 2PIC system using Opteon™ 2P50 across a wide range of conditions. The model aims to offer a comprehensive understanding of the system’s efficiency and potential improvements. The outcomes of this study are expected to contribute to the adoption of sustainable 2PIC cooling technologies in data centers.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Observations of wind farm wake recovery at an operating wind farm

Abstract. The interplay of momentum surrounding wind farms significantly influences wake recovery, affecting the speed at which wakes return to their freestream velocities. Under stable atmospheric conditions, wind farm wakes can extend over considerable distances, leading to sustained vertical momentum flux downstream, with variations observed throughout the diurnal cycle. Particularly in regions such as the US Great Plains, stable conditions can induce low-level jets (LLJs), impacting wind farm performance and power output. This study examines the implications of wake recovery using long-term observations of vertical momentum flux profiles across diverse atmospheric conditions. In these observations, several key findings were observed, such as (a) LLJ heights being altered downstream of a wind farm, especially when the LLJs are below 250 m above ground level; (b) a notable impact of LLJ height on wake recovery being observed using momentum flux profiles at upwind and downwind locations, wherein LLJs between 250 and 500 m above ground level resulted in larger momentum transfer within the wake (i.e., smaller velocity deficit) compared to LLJs below 250 m above ground level; (c) the largest momentum flux variability being observed during stable atmospheric conditions, with non-negligible variability observed during neutral and unstable atmospheric conditions; (d) detection of wake effects almost always being observed throughout the atmospheric boundary layer height; and finally (e) enhancement of wake recovery being observed in the presence of propagating gravity waves. These insights deepen our understanding of the intricate dynamics governing wake recovery in wind farms, advancing efforts to model and predict their behavior across varying atmospheric contexts. In addition, the performance of large-eddy-simulation-based semi-empirical internal boundary layer height model estimates incorporating real-world atmospheric and turbine inputs was evaluated using observations during LLJ conditions.

17 WIND ENERGY↗

Predictive numerical modeling of plasma-induced surface roughness and wettability evolution in LM-PAEK/CF tape

Plasma surface modification effectively enhances adhesion in thermoplastic composites, yet its impacts on high-performance polymers like low-melting polyaryletherketone (LM-PAEK) remain inadequately quantified. Here, this study integrates experimental analysis and numerical modeling to characterize surface roughness and wettability changes in LM-PAEK/carbon fiber composites treated with atmospheric plasma. Atomic Force Microscopy quantified surface topography (n = 10 per condition for contact angles), while static contact-angle assessments measured wettability. Roughness rapidly increased from ∼0.2 nm to 1.6 nm, and contact angle reduced from ∼90° to 24°, both stabilizing after 25–30 s of exposure. A semi-empirical, physics-informed framework was calibrated to these data, coupling surface chemistry via the Owens–Wendt decomposition with topography via the Wenzel roughness factor, and evaluated using out-of-sample (cross-validated) tests, while static contact angle assessments measured wettability. Numerical predictions matched experimental results closely (RMSE <5%, R 2 > 0.95). Incorporating material-specific parameters, the calibrated model supports plasma-treatment optimization and provides quantitative guidance for improving interfacial adhesion in thermoplastic composite manufacturing.

Atomic Force Microscopy↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Kβ X-ray Emission Spectra Analysis Using Bayesian Optimization

The Kβ X-ray emission spectrum of 3 d transition metals is rich with electronic and structural information due to strong exchange interactions with the valence shell of the metal, and has become crucial for understanding their spin and oxidation states. The spectrum is commonly treated using crystal-field multiplet theory, a semi-empirical theory that uses tunable parameters to control the strength of the effects present in X-ray emission spectroscopy (XES). However, determining the experimental values of these parameters remains a challenge. We present a methodology that applies Bayesian optimization to crystal-field multiplet theory to determine parameter values. The algorithm is tested on the X-ray emission spectra of a collection of Mn, Co, and Ni oxides. We are able to find optimal values for the four most impactful parameters: Slater−Condon reduction factors F dd , F pd , and G pd , and crystal field splitting 10 Dq . The algorithm produces significantly improved accuracy compared to current analysis methods, and probes interparameter dependencies by modeling the error landscape. This advancement enhances XES analysis by offering an approach of obtaining quantitative electronic structural information on 3 d transition metal valence shells, facilitating applications across various scientific fields.

Bayesian optimization↗