Search NASASearch

SEARCH · Search NASA

Results for “physics-based modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Regional Earthquake Ground Motion Simulations for Southern California With EQSIM: Insights From the 2008 Chino Hills, 2024 Highland Park, and 2021 Carson Earthquakes

This study presents physics-based, 3D simulations using the EQSIM framework for several earthquakes in the Los Angeles region. The primary objective was to assess the ability of deterministic physics-based ground motion simulations to reproduce the observed motions from historical events. The selected events included the mathematical equation M w 5.4 2008 Chino Hills, the mathematical equation M w 4.4 2024 Highland Park, and the mathematical equation M w 4.3 2021 Carson events. The simulated motions were evaluated by comparing the recorded and simulated seismograms, as well as the Fourier amplitude spectra, across multiple seismic stations. The SCEC 3D velocity model, CVM-S4.26.M01, was used to represent the regional geology, and ground motion simulations were carried out with a resolution of up to 5 Hz. The results indicate that the simulated motions captured the recorded motions up to approximately 4 Hz. While careful iterations regarding source parameters and corner frequencies were required, and, for the case of the Highland Park event, some of the near-source stations had relatively low accuracy, the present study established a positive step toward the utilization of physics-based simulations in practical applications. The computational efficiencies exhibited by EQSIM, especially on GPU clusters, further supported this assertion, as wall-clock times of simulations involving more than 10 billion grid points were as low as mathematical equation minutes. This permits ensemble simulations for a considered scenario event so that modeling uncertainties (e.g., source and geology) can be bracketed.

EQSIM

Hierarchical Reinforcement Learning of a Short-Range Bond-Order Potential for Silica: Analytic Embedding of Coordination with Classical Efficiency

Reinforcement learning (RL) has recently emerged as a data-efficient strategy to parametrize short-range interatomic potentials. Building on our past RL optimization of pairwise silica models, we extend the framework to a bond-order (Tersoff-type) potential that provides an analytic embedding of local coordination through a three-body term. A hierarchical RL workflow combining continuous-action Monte Carlo Tree Search and property-based rewards efficiently explores the 26-dimensional parameter space, sequentially optimizing lattice parameters, densities, angles, and cohesive energies of 21 silica polymorphs. The resulting models, Q-Tersoff and ML-Tersoff, reproduce the energetic ordering of low-energy phases and capture the angular correlations and amorphous structure factors of silica with improved fidelity over pairwise force fields, while remaining orders of magnitude faster than high-dimensional machine-learned potentials. Both models underperform for elastic constants and high-energy frameworks, delineating the limits of the current analytic form. The approach establishes a general and interpretable route to angle-aware, short-range potentials that bridge physics-based and machine-learned descriptions of silicate materials.

36 MATERIALS SCIENCE

Efficient Computation of Doppler-Broadened Elastic Scattering Kernel Moments Using Ladder-Operator Formulation

Anefficient routine for computing Legendre moments of the Doppler-broadened elastic scattering kernel, including resonance scattering effects, has been implemented in the ISOXML module of Griffin. Isotropic scattering in the center-of-mass system and the ideal gas model for target motion are assumed. A ladder-operator formulation is introduced to compute all Legendre moments from order 0 to N simultaneously, enabling near-linear scaling of computational cost with respect to the maximum Legendre order. A physics-based strategy for constructing outgoing energy grids has also been developed, in which a tailored base grid is combined with adaptive refinement to maintain accuracy while limiting the number of outgoing energy points. For energies between resonances, a constant cross-section model is employed to further reduce computational cost. In addition, a quantitative criterion is derived to determine isotope-wise cut-off incident energies based on a prescribed up-scattering probability coverage. For 238U, up to incident energies of approximately 75, 230, and 661 eV at 294, 900, and 2500 K (corresponding to a 2% up-scattering probability threshold), computation of P0 kernels requires 1–8 s and computation of P0–P5 kernels requires 0.4–4 min using a single thread, while maintaining 1–3% relative error in up-scattering probability. These results demonstrate that the proposed formulation enables accurate and computationally practical Doppler-broadened kernel generation for online multigroup cross-section production in Griffin.

Doppler-broadening

Scientific Discovery with Physics-Informed System Identification (Abbreviated Report)

My fellowship research focused on making physics-based simulations faster and more useful through machine learning. Many problems in science and engineering are governed by partial differential equations, but high-fidelity simulations are often too expensive to run repeatedly. I worked on improving Latent Space Dynamics Identification (LaSDI), a reduced-order modeling framework that compresses large simulation data sets into a smaller representation and then learns how that representation evolves over time. The motivation was to develop reduced models that remain accurate for more challenging systems, especially when predictions must remain reliable over long time intervals or when the underlying dynamics are more complicated than standard methods can easily handle. I also contributed to related work on Quandary, a high-performance software effort for simulation and control of open quantum systems, before focusing primarily on Latent Space Dynamics Identification methods. The main outcomes of the fellowship were two new algorithms (both of which were published), Rollout-LaSDI and Higher-Order LaSDI, together with supporting work on multi-stage Latent Space Dynamics Identification. Rollout-LaSDI improved long-term prediction by training the model to stay accurate over extended time horizons, and Higher-Order LaSDI broadened the method so it could model systems with higher-order time dynamics. My contributions to multistage Latent Space Dynamics Identification also helped show that its later training stages could be simplified without losing effectiveness, and that this behavior held across different model architectures and training strategies. Taken together, these advances improved the accuracy, flexibility, and practical value of reduced-order modeling tools for computational science.

97 MATHEMATICS AND COMPUTING

Seamlessly joining length scales: From atomistic thermal graphs to anisotropic continuum conductivity

Thermal transport in complex solids is governed by local structure, defects, and anisotropy, yet most continuum models still rely on oversimplified and homogenized conductivities. Here, we bridge atomistic and continuum descriptions by building finite element (FE) models directly from the site-projected thermal conductivity (SPTC), an atomic-level decomposition of the Green–Kubo thermal conductivity. We introduce a toolkit, the “Simulator Collection for Atomic-to-Continuum Scales (SCACS)”, which uses a graph neural network to predict SPTC on large atomic structures, coarse-grains these fields into anisotropic conductivity tensors, and embeds them into the heat-flow FE equation with a customized, anisotropy-aware adaptive mesh refinement scheme. Applied to silicon nanostructures, the resulting FE models act as representative volume elements, reproduce bulk conductivities, and capture interfacial and defect-driven anisotropy while maintaining thermodynamic consistency. Additionally, SCACS predicts experimental conductance trends and fields. This work demonstrates a general route for transferring atomistic transport information into device-scale thermal simulations with physics-based approximations.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

PPPL Report on Reduced Modeling of Fusion Alpha Transport in ARC Burning Plasmas

We are reporting on the modeling of fusion alpha particle transport in the planned ARC fusion device being designed by the CFS (Commonwealth Fusion Systems: https://cfs.energy). The ARC tokamak is designed to operate in a burning-plasma regime characterized by a substantial population of fusion-born alpha particles. Alfvén eigenmode (AE) stability is assessed both analytically and numerically, incorporating alpha-particle drive, ion Landau, and radiative damping from thermal species and collisional damping from trapped electrons. Regions of unstable and near-threshold AE activity are mapped across ARC’s operational parameter space. Linear stability analysis with NOVA indicates multiple, often marginally unstable AEs, extending to toroidal mode numbers up to n= 30. The present report focuses on the ARC flat-top operating point prior to the sawtooth event. Alpha-particle transport on timescales exceeding the neoclassical slowing-down time is assessed using the NUBEAM module [1][2] of the TRANSP code [3], employing transport coefficients derived from the RBQ quasilinear modeling (cf. Appendix B). These global simulations identify favorable and unfavorable operating regimes with respect to alpha confinement, pressure redistribution, and overall alpha-heating efficiency. We also evaluate additional transport mechanisms—including neoclassical tearing mode (TM)–induced stochasticity, sawtooth-driven redistribution, and toroidal-field ripple using the kick model (cf. Appendix C) which makes use of the guiding-center code ORBIT, see Section 5. The kick model is integrated into TRANSP to enable self-consistent predictions of alpha-driven current formation and sustainment within the ARC scenario. Sensitivity scans are performed over the mode frequency, rational-surface alignment, island width, mode amplitude, and proximity of the limiter to the plasma. Our study provides an initial, physics-based guidance for machine design, operational planning, and equilibrium control, ensuring adequate alpha confinement and robust self-heating performance in ARC. Our simulations mostly targeted worst case scenarios, e.g. for TMs and sawteeth. Overall, we expect benign effects for the ARC scenario investigated in this work on fusion alpha confinement and losses in the presence of AEs, tearing modes and sawteeth. This report addresses three thrusts identified at the outset. The first thrust focuses on analytic estimates of the parametric dependencies of EP relaxation based on local AE stability simulations (Section 3). The second thrust involves global evaluations of AE stability using the NOVA, RBQ, and NUBEAM codes (Section 4). Finally, we investigate alpha-particle transport driven by low-frequency instabilities associated with sawteeth and tearing modes (Section 5).

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

APOLLO: a facility-scale differentiable virtual accelerator for Fermilab

As the design complexity of modern accelerators grows, there is more interest in using advanced simulations that have fast execution time or yield additional insights like gradients. The FAST/IOTA facility has been working on implementing and experimentally validating an end-to-end digital twin that is both fast and gradient-aware, allowing for rapid prototyping of new software and experiments with minimal beam time costs. Our framework integrates physics and ML codes for linac and ring simulation through a set of generic interfaces between surrogate and physics-based sections. To reproduce device inputs and outputs, system state is exposed as a deterministic discrete event simulator. Because Fermilab is undergoing control system transition, both EPICS and ACNET frontends are supported. Recently, we have begun transitioning to a new community lattice standard, PALS, as well as developing standardized infrastructure for data ingest and normalization to prepare for model calibration during FAST proton injector commissioning. We discuss implementation details as well as challenges, and future plans to extend modelling to main complex proton accelerators like PIPII and Booster.

Kuklev, Nikita [Fermilab]

Low-Cost Sulfur Thermal Storage for Solar Industrial Process Heat Applications

Industrial process heat (IPH) is one of the largest energy demands in U.S., representing about 10% of all domestic energy consumption. Fuel costs to generate this industrial process heat are generally a top three cost for industry, a major component in American manufacturing competitiveness. Roughly 60% of US IPH demand (about 6,500 TBtu annually) falls in the medium-temperature range of 100–250 °C. While concentrated solar thermal (CST) technologies can provide a cost-effective source of heat in this temperature range, solar intermittency limits their adoption in industries that operate 24/7. Element 16 Technologies, Inc. developed a low-cost sulfur thermal energy storage (TES) technology to bridge this gap by capturing excess solar heat during the day and dispatching it reliably during non-solar hours. The core innovation is the use of sulfur, an abundant, industrial waste byproduct that costs ten times less than molten salt used in commercial TES systems. The overall goal of the project was to advance the design and development of molten sulfur TES to a manufacturing-relevant prototype stage for solar industrial process heat applications, while establishing and validating a realistic pathway to commercial success. Key tasks included corrosion and mechanical durability testing to identify cost-effective materials, design investigations using physics-based simulation tools, techno-economic evaluations of system lifetime costs, and pilot-scale testing for performance verification. Corrosion testing of steel alloys under cyclic molten sulfur conditions showed that austenitic stainless steels in the 300 series performed particularly well, with no structural degradation of welds or joints. Thermal cyclic testing of pilot sulfur TES units up to 1.5 MWh quantified charge/discharge rates, heat losses, round-trip efficiency and validated the system's capability to operate effectively under intermittent charging conditions. A techno-economic model, informed by sulfur TES performance model validated using pilot test data, showed that hybrid solar+sulfur TES+NG boiler systems are economically competitive with incumbent natural gas boilers for multiple locations in the southwest US. In summary, this project established molten sulfur TES as a technically viable pathway to improve economic competitiveness of American manufacturing by lowering the cost of solar industrial process heat.

14 SOLAR ENERGY

Crowdsourcing the Frontier: Advancing Hybrid Physics‐ML Climate Simulation via a $\$$50,000 Kaggle Competition

Subgrid machine-learning (machine learning [ML]) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and ML researchers opened up the offline aspect of this problem to the broader ML and data science community with the release of ClimSim, a NeurIPS Data sets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art results on certain metrics such as zonal mean bias patterns and global Root Mean Squared Error, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

Environmental sciences

Plume Impingement Software Module for Real-Time Proximity Operations

Successfully executing proximity operations in space, such as docking or in-orbit servicing, requires sophisticated spacecraft design that accounts for induced environments. As a chaser vehicle’s attitude control thrusters fire, they create rarefied plumes that can impact the target vehicle, with the potential to overload components, exceed thermal limits, and spin the target vehicle out of control. High-fidelity simulations of the thruster plume impingement environment require the direct simulation Monte Carlo (DSMC) method, but DSMC is too computationally expensive to simulate proximity operations that involve thousands of thruster firings. For this analysis to be tractable, engineering models of the plume flowfield and impingement events are used to simulate these trajectories [1]. Currently, on-orbit plume impingement environments are modeled through an inefficient open-loop analysis cycle where the vehicle’s flight controller and plume impingement teams iterate on the trajectories until they pass the target vehicle’s plume requirements. As complex on-orbit missions evolve and become more frequent, lengthy design cycles will become operational bottlenecks. To address this gap, this work develops an advanced plume impingement module capable of operating at real-time scale that can be integrated with existing mission planning tools and onboard flight systems. The plume module leverages state-of-the-art plume simulation techniques [2] to deliver fast, physics-based impingement predictions in a software architecture that can be tailored to diverse proximity operations scenarios. A prototype of this plume impingement module is built to demonstrate the feasibility of real-time performance. This prototype completes plume impingement calculations in microseconds per target geometry mesh point. The software serves as a foundational capability for plume-aware trajectory design, operational risk assessment, and future autonomous decision-making systems.

Plume Impingement

Draft Impacts: Modeling and Lab Validation (CRADA Final Report)

Chimney draft is a significant source of variability in real-world wood heater performance, yet laboratory tests conducted for certification do not capture this variability. Because draft varies with climate, chimney height, and home conditions, a heater that performs well in the lab can perform quite differently when installed in a home. Until now, wood heater manufacturers and installers have lacked the tools needed to anticipate how draft will vary across real installations or to advise on corrective measures when draft is too low or too high. The goal of this project was to develop and validate an open-source draft prediction tool for cordwood heater chimney systems. Over 18 months, Lawrence Berkeley National Laboratory (LBNL) and the Hearth, Patio & Barbecue Association (HPBA) collaborated to build a physics-based, Python tool that predicts how chimney draft evolves during operation, from ignition through steady burning. HPBA convened a stakeholder group of manufacturers and venting experts who provided feedback throughout development. LBNL validated the tool against laboratory measurements from a catalytic and a non-catalytic cordwood heater operated across a range of chimney heights and room-pressure conditions.

42 ENGINEERING

Auxiliary heating and current drive physics for the ST-E1 fusion power plant

This work describes the physics basis for the proposed auxiliary heating and current drive system on the ST-E1 fusion power plant. The ST-E1 flattop plasma considered here is fully non-inductive with a bootstrap fraction of 0.9 and the remaining current driven by EC waves. Using the recently published physics-based optimization method for EC launchers (Lopez et al 2025 Plasma Phys. Control. Fusion 67 055012), we show that the target flattop ECCD can be achieved with a net efficiency of 52 kA MW −1 using fundamental O-mode (O1) with frequency range 160–200 GHz launched from the low-field side top half of the vacuum vessel (LFS top-launch). From considering two candidate rampup scenarios, we conclude that LFS top-launch O1 ECCD can be equally effective during the early stages of plasma operation, although poloidal steering might be needed. X-mode waves injected from the LFS midplane are also shown to be effective for rampup even when T e < 1 keV. We also present modeling results for the pre-conceptual design of an ICRH system proposed for ST-E1. Using TORIC, we find that an ICRH system aiming for 42–48 MHz and toroidal mode number n φ ~ 10 robustly achieves dominant ion damping via Helium-3 minority heating transitioning to second-harmonic Tritium heating. We then show that such waves can be efficiently generated by a 5-strap traveling-wave antenna (TWA) using the Petra-M code. The TWA has a 40–45 MHz passband within which ~60% of the power entering the TWA is coupled to the plasma with the remaining ~40% of the power being transmitted through the TWA and possibly recirculated; the power reflected back into the transmission lines is negligible. This passband structure persists even when the evanescent distance is increased by a factor of two, or when the magnetic-field angle is increased by 30°, demonstrating inherent load resilience that will be crucial for effective ICRH on ST-E1.

electron cyclotron

Explainable physics-based constraints on reinforcement learning for accelerator optimization

We present a reinforcement learning (RL) framework for optimizing particle accelerator experiments that builds explainable physics-based constraints on agent behavior. The goal is to increase transparency and trust by letting users verify that the agent’s decision-making process incorporates suitable physics. Our algorithm uses a learnable surrogate function for physical observables, such as energy, and uses them to fine-tune how actions are chosen. This surrogate can be represented by a neural network or by an interpretable sparse dictionary model. We test our algorithm on a range of particle accelerator optimization environments designed to emulate the Continuous Electron Beam Accelerator Facility at Jefferson Lab. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment. In addition, we find that the introduction of a physics-based surrogate enables our RL algorithms to reliably converge for difficult high-dimensional accelerator optimization environments.

explainability

Optimizing district energy systems by integrating Borehole Thermal Energy Storage Using a Mixed-Integer Linear Programming g-function framework with a Multi-Timescale Rolling Horizon method

Shallow geothermal has gained increasing attention in recent years; however, a reliable framework for its accurate incorporation into large-scale energy system optimization remains lacking. This study proposes a Mixed-Integer Linear Programming (MILP) framework combined with the g-function approach to integrate Borehole Thermal Energy Storage (BTES) technology into energy system optimization. Validation against a Modelica-based reservoir network simulation demonstrates that the proposed framework effectively captures the ground thermal response under varying energy loads and accurately estimates the borefield energy supply. To enhance scalability, a Rolling Horizon with Multi-Timescale (RH-MTS) method is further introduced, reducing computational time by 73 % for the 1-year optimization model with only minor loss of optimality. The framework is demonstrated through the case study of the UC Berkeley campus. Results indicate that BTES is a cost-effective and low-carbon solution: two borefields comprising 382 boreholes can meet 8.0 % and 6.6 % of the total campus heating and cooling demand, respectively, at an average energy rate of 0.70–0.77 USD/kWh and carbon intensity of 0.54 kg-CO2/kWh. Short-term analysis reveals a 35%–65% decline in BTES energy flow after 3–6 months of continuous heating/cooling operation, while long-term simulation shows that annual energy production of BTES can vary by up to 12.0 % after four years before stabilizing. Overall, this study develops a novel optimization framework that couples physics-based g-function method with MILP optimization framework, thereby advancing methodological development for shallow-geothermal integration and providing actionable guidance for BTES deployment in district-energy systems.

Yang, Jiahui

APOLLO: a facility-scale differentiable virtual accelerator at Fermilab FAST/IOTA

As the design complexity of modern accelerators grows, there is more interest in using advanced simulations that have fast execution time or yield additional insights like gradients. The FAST/IOTA facility has been working on implementing and experimentally validating an end-to-end digital twin that is both fast and gradient-aware, allowing for rapid prototyping of new software and experiments with minimal beam time costs. Our framework integrates physics and ML codes for linac and ring simulation through a set of generic interfaces between surrogate and physics-based sections. To reproduce device inputs and outputs, system state is exposed as a deterministic event loop in a specialized discrete event simulator architecture. Because Fermilab is undergoing control system transition, several APIs were implemented as final user interfaces - a fully asynchronous EPICS soft IOC, a gRPC-based Data Pool Manager (DPM), and legacy ACNET protocols. We discuss implementation details as well as challenges handling live data assimilation and future plans to extend modelling to main complex proton accelerators like PIPII and Booster.

Kuklev, Nikita [Fermilab]

A Computational Framework to design 3D stiffness gradient acoustic metamaterials for impedance matching

Acoustic waves play a crucial role in various applications, including medical imaging, non-destructive testing, and sonar systems. One of the significant challenges in these applications is impedance matching, which is essential for minimizing reflections and maximizing the transfer of acoustic energy between different media. Acoustic metamaterials offer a promising solution to this challenge. In addition to impedance control, gradient stiffness can enhance structural efficiency and enable spatial control of wave propagation, making it a valuable feature in acoustic metamaterial design. In this pa- per, we present our developed computational method to design 3D stiffness gradient acoustic metamaterials for impedance matching. The key steps in our approach include generating initial designs using a periodic covariance function to provide unit cells that are both periodic on the boundaries and randomly formed inside the unit cell. Furthermore, we integrated manufacturing constraints into the design process, ensuring that the structures are interconnected for fabrication. We propose two computational optimization algorithms: GenUnit, based on a non-dominated sorting genetic algorithm (NSGA-II), and MLMatch, which leverages differentiable machine learning. The two approaches are not separate contributions but complementary com- ponents of a unified framework. GenUnit requires no training data and directly interfaces with physics-based simulations, making it highly accurate but slower for large-scale exploration. In contrast, MLMatch is data-hungry during training but, once trained, enables near-instantaneous inference and broad design-space coverage. Together, they form a hybrid strategy: ML- Match rapidly explores the global design space, and GenUnit provides local refinement with high-fidelity accuracy. This balance between training cost, inference time, and precision is the motivation for including both methods in the same study. We applied this dual-algorithm framework to generate two metallic-based metamaterial designs that match the acoustic impedance of water while exhibiting a controlled gradient in stiffness (from stiff to soft). The stiffness gradient is particularly advantageous in applications where one side of the structure must interface with soft or sensitive surfaces, such as human tissue or delicate components. Here, this work paves the way for improved materials in various acoustic applications, particularly in ultrasound devices, by providing better impedance.

Metamaterial

Differentiable multiphase flow model for physics-informed machine learning in reservoir pressure management

Accurate subsurface reservoir pressure control is extremely challenging due to geological heterogeneity and multiphase fluid-flow dynamics. Predicting behavior in this setting relies on high-fidelity physics-based simulations that are computationally expensive. Yet, the uncertain, heterogeneous properties that control these flows make it necessary to perform many of these expensive simulations, which is often prohibitive. To address these challenges, we introduce a physics-informed machine learning workflow that couples a fully differentiable multiphase flow simulator, which is implemented in the DPFEHM framework with a convolutional neural network (CNN). The CNN learns to predict fluid extraction rates from heterogeneous permeability fields to enforce pressure limits at critical reservoir locations. By incorporating transient multiphase flow physics into the training process, our method enables more practical and accurate predictions for realistic injection-extraction scenarios compared to previous works. To speed up training, we pretrain the model on single-phase, steady-state simulations and then finetune it on full multiphase scenarios, which dramatically reduces the computational cost. We demonstrate that high-accuracy training can be achieved with fewer than three thousand full-physics multiphase flow simulations – compared to previous estimates requiring up to ten million. This drastic reduction in the number of simulations is achieved by leveraging transfer learning from much less expensive single phase simulations.

25 ENERGY STORAGE

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING