Search NASASearch

SEARCH · Search NASA

Results for “machine learning potential”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

When more data hurts: Optimizing data coverage while mitigating diversity-induced underfitting in an ultrafast machine-learned potential

Machine-learned interatomic potentials (MLIPs) are becoming an essential tool in materials modeling. However, optimizing the generation of training data used to parametrize the MLIPs remains a significant challenge. This is because MLIPs can fail when encountering local environments too different from those present in the training data. The difficulty of determining a priori the environments that will be encountered during molecular dynamics simulation necessitates diverse, high-quality training data. Here, this study investigates how training data diversity affects the performance of MLIPs using the Ultra-Fast force field (UF 3 ) to model amorphous silicon nitride. We employ expert and autonomously generated data to create the training data and fit four force field variants to subsets of the data. Our findings reveal a critical balance in training data diversity: insufficient diversity hinders generalization, while excessive diversity can exceed the MLIP's learning capacity, reducing simulation accuracy. Specifically, we found that the UF 3 variant trained on a subset of the training data, in which nitrogen-rich structures were removed, offered vastly better prediction and simulation accuracy than any other variant. By comparing these UF 3 variants, we highlight the nuanced requirements for creating accurate MLIPs, emphasizing the importance of application-specific training data to achieve optimal performance in modeling complex material behaviors.

ab initio molecular dynamics

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE

Systematic softening in universal machine learning interatomic potentials

Machine learning interatomic potentials (MLIPs) have introduced a new paradigm for atomic simulations. Recent advancements have led to universal MLIPs (uMLIPs) that are pre-trained on diverse datasets, providing opportunities for universal force fields and foundational machine learning models. However, their performance in extrapolating to out-of-distribution complex atomic environments remains unclear. In this study, we highlight a consistent potential energy surface (PES) softening effect in three uMLIPs: M3GNet, CHGNet, and MACE-MP-0, which is characterized by energy and force underprediction in atomic-modeling benchmarks including surfaces, defects, solid-solution energetics, ion migration barriers, phonon vibration modes, and general high-energy states. The PES softening behavior originates primarily from the systematically underpredicted PES curvature, which derives from the biased sampling of near-equilibrium atomic arrangements in uMLIP pre-training datasets. Our findings suggest that a considerable fraction of uMLIP errors are highly systematic, and can therefore be efficiently corrected. We argue for the importance of a comprehensive materials dataset with improved PES sampling for next-generation foundational MLIPs.

36 MATERIALS SCIENCE

Teacher-student training improves the accuracy and efficiency of machine learning interatomic potentials

Machine learning interatomic potentials (MLIPs) are revolutionizing the field of molecular dynamics (MD) simulations. Recent MLIPs have tended towards more complex architectures trained on larger datasets. The resulting increase in computational and memory costs may prohibit the application of these MLIPs to perform large-scale MD simulations. Herein, we present a teacher-student training framework in which the latent knowledge from the teacher (atomic energies) is used to augment the students' training. We show that the light-weight student MLIPs have faster MD speeds at a fraction of the memory footprint compared to the teacher models. Remarkably, the student models can even surpass the accuracy of the teachers, even though both are trained on the same quantum chemistry dataset. Our work highlights a practical method for MLIPs to reduce the resources required for large-scale MD simulations.

36 MATERIALS SCIENCE

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Computational investigation of water glasses using machine-learning potentials

The molecular origins of water’s anomalous properties have long been a subject of scientific inquiry. The liquid–liquid phase transition hypothesis, which posits the existence of distinct low-density and high-density liquid states separated by a first-order phase transition terminating at a critical point, has gained increasing experimental and computational support and offers a thermodynamically consistent framework for many of water’s anomalies. However, experimental challenges in avoiding crystallization near the postulated liquid–liquid critical point have focused attention to water’s canonical glassy states: low-density and high-density amorphous ice. Here, we use two Deep Potential machine-learning models, trained on the Strongly Constrained and Appropriately Normed density functional and the highly accurate Many-Body Polarizable potential, to conduct an investigation of water’s glassy phenomenology based on quantum mechanical calculations. Despite not being explicitly trained on amorphous ices, both models accurately capture the structure and transformation of the water glasses, including their interconversion along different thermodynamic paths. Isobaric quenching of liquid water at various pressures generates a continuum of intermediate amorphous ices and density fluctuations increase near the liquid–liquid critical pressure. The glass transition temperatures of the amorphous ices produced at different pressures exhibit two distinct branches, corresponding to low-density and high-density amorphous ice behaviors, consistent with experiment and the liquid–liquid transition hypothesis. Extrapolating transformation pressures from isothermal compressions to experimental compression rates brings our simulations into excellent agreement with data. Our findings demonstrate that machine-learning potentials trained on equilibrium phases can effectively model nonequilibrium glassy behavior and pave the way for studying long-timescale, out-of-equilibrium processes with quantum mechanical accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Microscopic insights into the solvation of polyethylene glycol chains in water: A machine learning potential approach

Polyethylene glycol (PEG) is a structurally simple, nontoxic, and water-soluble polymer widely utilized in medical and pharmaceutical applications. Notably, when a PEG chain is immersed in water, the surrounding water molecules play a key role in driving conformational changes of this macromolecule. In this study, we explore the solvation behavior of PEG under mechanical strain using molecular dynamics simulations, with an interatomic potential obtained from machine learning. Our focus is on the transition from the favored coil-like conformation to an extended one under external force. Through analyses of radial distribution functions, hydrogen bonding, and solvation dynamics, we uncover how mechanical stretching influences the local hydration environment. Furthermore, we disentangle the enthalpic and entropic contributions to the conformational stability of PEG in water. Surprisingly, our neural network potential model identifies dewetting of PEG C-atoms, and not water H-bonding with PEG O-atoms, as the main enthalpic driving force for the coiling of PEG in water.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Atomistic Simulations of Thermal and Chemical Expansions of PrNi x Co 1‐x O 3‐δ Accelerated by Machine Learning Potentials

The electrodes and solid-state electrolytes in protonic ceramic electrochemical cells (PCECs) experience significant lattice expansions when exposed to high steam concentrations at elevated temperatures. In this paper, phonon calculations based on a new machine learning potential (MLP) are employed to elucidate the volume expansions of the proton-conducting PrNi x Co 1-x O 3-δ (PNC) lattices, manifested under a combined influence of oxygen vacancies (V$^{\cdot\cdot}_O$ ) and proton uptake (OH$^{\cdot}_O$ ) in the bulk at varying Ni/Co occupancies. It is revealed that the Ni/Co occupancy contributes to thermal and chemical expansions differently, where thermal expansions are related to Co occupancy. In contrast, chemical expansions are more closely associated with the Ni occupancy. Both V$^{\cdot\cdot}_O$ and OH$^{\cdot}_O$ lead to higher thermal expansions when compared to the pristine PNC. The temperature increase will negatively impact the hydration-induced chemical expansions. For combined thermal and chemical expansions, it is predicted that the strategies that boost the PCEC's electrochemical performance may harm the electrode–electrolyte interfacial stability, when the Ni occupancy is high, due to severe chemical expansions. Mitigating chemical expansions of the Ni-abundant PNC will benefit the interfacial stability. Finally, the presented computational methods for phonon calculations, based on emerging machine learning interatomic potential techniques are anticipated to have a lasting impact on future PCEC development.

computational prediction

Accelerating the Structure Exploration of Diverse Bi–Pt Nanoclusters via Physics‐Informed Machine Learning Potential and Particle Swarm Optimization

Bimetallic Bi–Pt nanoclusters exhibit diverse structural motifs, including core-shell, Janus, and mixed alloy configurations, due to the unique bonding characteristics between Bi and Pt atoms. Using density functional theory refinements from ChIMES physically machine-learned potential and CALYPSO particle swarm optimization global searches, 34 Bi20-Pt20 nanoclusters are systematically classified. The results reveal that Bi atoms predominantly occupy surface sites, driven by charge transfer effects. Cohesive energy trends alone prove insufficient for structure differentiation, necessitating a data-driven approach employing principal component analysis and K-means clustering. Furthermore, vibrational, electronic, and infrared spectral analyses provide additional insights into structure-property relationships. The findings offer an original framework for the automated classification and analysis of bimetallic nanoclusters, enhancing the understanding of their stability and functional properties.

bimetallic nanoparticles

Generalizable machine learning potentials for quantum-accurate predictions of non-equilibrium behavior in 2D materials

Machine learning interatomic potentials (ML-IAPs) are emerging as transformative tools in materials modeling, promising quantum-level accuracy at a fraction of the computational cost. However, their ability to generalize beyond equilibrium configurations and to reliably capture defect- and temperature-driven behavior remains underexplored. Here, we develop and benchmark two state-of-the-art ML-IAPs, Spectral Neighbor Analysis Potential (SNAP) and Allegro, on a comprehensive dataset for monolayer MoSe₂. Using density functional theory (DFT) as the reference, we evaluate their performance in capturing stress–strain behavior, phase transition energetics, defect evolution, edge stability, and fracture toughness. Allegro, a deep equivariant neural network potential, surpasses both SNAP and the classical Tersoff potential in accuracy, efficiency, and transferability. Importantly, both ML potentials accurately reproduce experimental fracture measurements and ab initio predictions of inversion domain formation—phenomena well beyond their training sets. Our findings establish ML-IAPs as viable replacements for traditional force fields in the study of non-equilibrium mechanical phenomena, enabling large-scale, high-fidelity simulations in 2D materials and beyond. In conclusion, this work provides a broadly applicable framework for data-driven modeling of structural and functional transformations under extreme conditions.

2D materials

Machine learned potential for high-throughput phonon calculations of metal—organic frameworks

Metal–organic frameworks (MOFs) are highly porous and versatile materials studied extensively for applications such as carbon capture and water harvesting. However, computing phonon-mediated properties in MOFs, like thermal expansion and mechanical stability, remains challenging due to the large number of atoms per unit cell, making traditional Density Functional Theory (DFT) methods impractical for high-throughput screening. Recent advances in machine learning potentials have led to foundation atomistic models, such as MACE-MP-0, that accurately predict equilibrium structures but struggle with phonon properties of MOFs. In this work, we developed a workflow for computing phonons in MOFs within the quasi-harmonic approximation with a fine-tuned MACE model, MACE-MP-MOF0. The model was trained on a curated dataset of 127 representative and diverse MOFs. The fine-tuned MACE-MP-MOF0 improves the accuracy of phonon density of states and corrects the imaginary phonon modes of MACE-MP-0, enabling high-throughput phonon calculations with state-of-the-art precision. The model successfully predicts thermal expansion and bulk moduli in agreement with DFT and experimental data for several well-known MOFs. These results highlight the potential of MACE-MP-MOF0 in guiding MOF design for applications in energy storage and thermoelectrics.

Elena, Alin Marin

Liquid-liquid phase transition of hydrogen and its critical point: Analysis from ab initio simulation and a machine-learned potential

We simulate high-pressure hydrogen in its liquid phase close to molecular dissociation using a machine-learned interatomic potential. The model is trained with density functional theory (DFT) forces and energies, with the Perdew-Burke-Ernzerhof (PBE) exchange-correlation functional. We show that an accurate NequIP model, an E(3)-equivariant neural network potential, accurately reproduces the phase transition present in PBE. Moreover, the computational efficiency of this model allows for substantially longer molecular dynamics trajectories, enabling us to perform a finite-size scaling (FSS) analysis to distinguish between a crossover and a true first-order phase transition. Here, we locate the critical point of this transition, the liquid-liquid phase transition (LLPT), at 1200-1300 K and 155-160 GPa, a temperature lower than most previous estimates and close to the melting transition.

08 HYDROGEN

Comparison of DeePMD, MTP, GAP, ACE and MACE Machine‐Learned Potentials for Radiation‐Damage Simulations: A User Perspective

Accurate and efficient interatomic potentials are essential for molecular dynamics (MD) simulations of radiation damage, gas diffusion, and phase stability in complex ceramics such as LiAlO 2 , especially under extreme conditions relevant to tritium production. Here, we evaluate the performance of six machine-learned interatomic potentials (MLIPs), moment tensor potential (MTP), Gaussian approximation potential, deep potential (DeePMD), atomic cluster expansion (ACE), message-passing ACE (multilayer atomic cluster expansion (MACE) pretrained) and MACE (trained from-scratch), all trained on the same density functional theory dataset with inclusion of tritium. The MLIPs are benchmarked against traditional Buckingham and ReaxFF potentials in terms of energy accuracy, density predictions, thermal equilibration behavior, threshold displacement energy (E d ), tritium diffusivity, and computational cost. Among the models, MTP shows the best overall balance between efficiency and accuracy, with low force and energy errors and realistic E d values for Li and Al. The ACE and MACE (pretrained and trained from scratch) models exhibit high E d (>200 eV) and unphysical pair interactions. DeePMD underestimates Ed due to overly repulsive behavior even at equilibrium distances. All models over-estimate tritium diffusion but the pretrained MACE model behaves well during tritium-diffusion simulations up to 500 K, maintaining diffusivities in the physically consistent 10 −11 m 2 /s range. Finally, we quantify the computational cost of each potential in large-scale atomic/molecular massively parallel simulator, finding that only MTP is more efficient than traditional empirical potentials, while others are significantly more expensive. These findings explain the trade-offs between accuracy and computational cost in MLIP development and provide essential guidance for use in high-throughput radiation damage and gas diffusion simulations in nuclear ceramics.

74 ATOMIC AND MOLECULAR PHYSICS

Weighted active space protocol for multireference machine-learned potentials

Multireference methods such as multiconfiguration pair-density functional theory accurately capture electronic correlation in systems with strong multiconfigurational character, but their cost precludes direct use in molecular dynamics. Combining these methods with machine-learned interatomic potentials (MLPs) can extend their reach. However, the sensitivity of multireference calculations to the choice of the active space complicates the consistent evaluation of energies and gradients across structurally diverse nuclear configurations. To overcome this limitation, we introduce the weighted active space protocol (WASP), a systematic approach to assign a consistent active space for a given system across uncorrelated configurations. By integrating WASP with MLPs and enhanced sampling techniques, we propose a data-efficient active learning cycle that enables the training of an MLP on multireference data. We demonstrated the approach on the TiC + -catalyzed C–H activation of methane, a reaction that poses challenges for Kohn–Sham density functional theory due to its significant multireference character. This framework enables accurate and efficient modeling of catalytic dynamics, establishing a paradigm for simulating complex reactive processes beyond the limits of conventional electronic-structure methods.

enhanced sampling

Designing a quantum-accurate machine-learning potential to enable large-scale simulations of deuterium under shock

Large-scale molecular dynamics of deuterium under shock can elucidate kinetic processes vital to the target design in inertial confinement fusion and high-energy-density experiments. However, modeling the complex evolution of this material from an insulating molecular state at ambient pressure to an ionized, atomic fluid under strong shock is beyond the capability of simple pair and even bond order potentials. We thus train a quantum-accurate and broadly transferable machine-learning interatomic potential for deuterium using the Chebyshev Interaction Model for Efficient Simulations framework. We show that due to an improved description of the molecular-to-atomic transition, our model is able to better reproduce the ab initio equation of state, radial distribution functions, and principal Hugoniot than bond order potentials. This represents an important step toward large-scale quantum-accurate and nonequilibrium simulations of complicated systems under dynamic changes including phase transitions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Nature of molybdenum carbide surfaces for catalytic hydrogen dissociation using machine-learned potentials: an ensemble-averaged perspective

Molybdenum carbides with an electronic structure similar to noble metals have gained attention as a promising low-cost catalyst for biomass valorization and the hydrogen evolution reaction. However, our fundamental understanding of the catalyst surface and how different phases of these catalysts behave at varying reaction conditions is limited to ground state density functional theory calculations as ab initio molecular dynamics (AIMD) is computationally prohibitive at relevant length and time scales. Here, in this work, we train a multi-atomic cluster expansion (MACE) machine-learned interatomic potentials (MLIP) to study hydrogen dissociation and dynamics over Mo, δ-MoC, α-Mo 2 C, and β-Mo 2 C surfaces at varying temperatures and hydrogen partial pressures. Our simulations identify unique and different molecular and atomic hydrogen adsorption sites on different surfaces that do not depend on the temperature. At low hydrogen pressures, the surface coverage is monolayer, which transitions to two-layer adsorption at higher pressures. We find that atomic hydrogen diffusion and recombinations are preferred over molybdenum atom hollow sites, while the diffusion over carbon-terminated facets was negligible, signifying particularly strong C–H interactions. In contrast, molecular hydrogen adsorption occurs mostly atop Mo or the bridging sites. At a comparable hydrogen loading, β-Mo 2 C (001) is the most active surface for hydrogen dissociation reaction. This work provides insights into the dynamic nature of the hydrogen dissociation chemistry and the diversity of hydrogen adsorption sites on molybdenum carbides.

08 HYDROGEN

Improving Bond Dissociations of Reactive Machine Learning Potentials through Physics-Constrained Data Augmentation

In the field of computational chemistry, predicting bond dissociation energies (BDEs) presents well-known challenges, particularly due to the multireference character of reactive systems. Many chemical reactions involve configurations where single-reference methods fall short, as the electronic structure can significantly change during bond breaking. As generating training data for partially broken bonds is a challenging task, even state-of-the-art reactive machine learning interatomic potentials (MLIPs) often fail to predict reliable BDEs and smooth dissociation curves. By contrast, simple and inexpensive physics-based models, such as the well-established Morse potential, do not suffer from any such limitations. This work leverages the Morse potential to improve reactive MLIPs by augmenting the training data set with inexpensive Morse data along the dissociation pathways. Further, this physics-constrained data augmentation (PCDA) approach results in MLIPs with smooth bond dissociation curves as well as near coupled-cluster level BDEs, all without requiring any expensive multireference quantum mechanical calculations. A case study for methane combustion demonstrates how the PCDA approach can improve an existing reactive MLIP, namely, ANI-1xnr. In conclusion, not only are the BDEs and bond dissociation curves for all radicals and molecules significantly improved compared to ANI-1xnr but the PCDA-trained MLIP retains the reliability of ANI-1xnr when performing reactive molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Modeling Equilibrium Solid–Liquid Interfaces under Effective Constant Chemical Potential Using Machine Learning Interatomic Potentials

The chemical potential (μ) of species in solution is essential for understanding various chemical processes at interfaces. Molecular dynamics (MD) simulations, constrained by fixed compositions, cannot maintain constant chemical potential with reference to a targeted concentration or chemical potential under nonequilibrium or dynamic conditions, as solute species can migrate to the interface and deplete (or enrich) the bulk due to solute-interface interactions. In this study, we introduce a simple and computationally efficient approach named iterative quasi-constant chemical potential molecular dynamics (iqCμMD) simulation, which helps simulate targeted molar concentrations of species in solution. iqCμMD overcomes the limitations of conventional MD by adjusting the number of species in the solution to reach a target bulk concentration (chemical potential), which allows simulation of the interface under the bulk conditions comparable to experiment. We demonstrate our approach using machine learning interatomic potential (MLIP)-based MD simulations of the Na 2 SO 4,aq –graphene interface, and to show the transferability of our approach, we also perform classical force field-based MD simulations of NaCl aq –air and NaCl aq –graphite interfaces, which produce comparable results to previous CμMD simulations. Our results also show that the iqCμMD approach efficiently achieves the desired bulk ion concentration within two iterations, and by utilizing MLIPs, we can achieve converged results using relatively small-scale simulations compared to previous CμMD simulations. By combining iqCμMD with MLIP-driven simulations, solid–liquid interfaces can be modeled under an effective constant chemical potential with DFT-level accuracy. Here, we show that iqCμMD offers a robust and simple computational framework for constant chemical potential simulations, as its only requirement is to be able to converge interfacial simulations with a measurable bulk region.

Chemical structure