Search NASA⌕ Search

SEARCH · Search NASA

Results for “training data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Machine Learning-Driven Solvent Screening for Biobased 2,3-Butanediol Extraction

Biobased 2,3-butanediol (2,3-BDO) is a valuable biomass-derived chemical due to its versatility in being transformed into a wide variety of products. However, the separation and purification of 2,3-BDO from fermentation broth remain a significant challenge owing to its high boiling point and hydrophilic nature. Herein, we developed a machine learning (ML)-based screening workflow that uses molecular calculations as training data and requires only a small number of experimental measurements for validation to identify alternative solvent candidates for the liquid–liquid extraction (LLE) of 2,3-BDO from aqueous solution. In particular, 130 density functional theory (DFT) calculations with the implicit solvation method not only built a correlation between the computational partition coefficient and the experimental distribution coefficient of 2,3-BDO but also parameterized an Extra-Trees ML model to screen the distribution coefficient for a wider range of 6717 organic solvents. The experimental measurements of only 24 solvents were needed to validate the computational results. A list of 50 prioritized solvents was proposed for 2,3-BDO LLE, and seven additional experimental measurements were conducted to further verify our selected solvents. The impact of the extraction temperature and solvent-to-feed ratio was also investigated for selected solvents in experiments. Furthermore, this work suggested alternative solvents for 2,3-BDO LLE and proposed a versatile workflow that requires fewer experiments and can be applied to a broader range of LLE studies.

Extraction↗

Comparative Analysis of TCR and TCR-pMHC Complex Structure Prediction Tools

The rapid development of computational approaches for predicting the structures of T cell receptors (TCRs) and TCR-peptide-major histocompatibility (TCR-pMHC) complexes, accelerated by AI breakthroughs such as AlphaFold, has made it feasible to calculate these structures with increasing accuracy. Although these tools show great potential, their relative accuracy and limitations remain unclear due to the lack of standardized benchmarks. Here, we systematically evaluate seven tools for predicting isolated TCR structures together with six tools for predicting TCR-pMHC complex structures. The methods include homology-based approaches, general prediction tools using AlphaFold, TCR-specific tools derived from AlphaFold2, and the newly developed tFold-TCR model. The evaluation uses a post-training data set comprising 40 αβ TCRs and 27 TCR-pMHC complexes (21 Class I and 6 Class II). Model accuracy is assessed at global, local, and interface levels using a variety of metrics. We find that each tool offers distinct advantages in various aspects of its predictions. AlphaFold2, AlphaFold3, and tFold-TCR excel in overall accuracy of TCR structure prediction, and TCRmodel2 and AlphaFold2 perform well in overall accuracy of TCR-pMHC structure prediction. However, TCR-specific tools derived from AlphaFold2 show lower accuracy in the framework region than both homology-based methods and general-purpose tools such as AlphaFold, and challenges remain for all in modeling CDR3 loops, docking orientations, TCR-peptide interfaces, and Class II MHC-peptide interfaces. Furthermore, these findings will guide researchers in selecting appropriate tools, emphasize the importance of using multiple evaluation metrics to assess model performance, and offer suggestions for improving TCR and TCR-pMHC structure prediction tools.

Chemical structure↗

Machine Learning a Simple Interpretable Short-Range Potential for Silica

A wide array of models, spanning from computationally expensive ab initio methods to a spectrum of force-field approaches, have been developed and employed to probe silica polymorphs and understand growth processes and atomic-level dynamical transitions in silica. However, the quest for a model capable of making accurate predictions with high computational efficiency for various silica polymorphs is still ongoing. Recent developments in short-range machine-learned models, such as GAP and NNPScan, have shown promise in providing reasonable descriptions of silica, but their computational cost remains high compared to force fields such as BKS which are based on simple interpretable functional forms. Here, in this study, we build on the recent success of our reinforcement learning (RL) workflow to derive a new set of optimal parameters for a promising short-range BKS-based model proposed by Soules. We use RL to navigate the eight-dimensional parameter space of the Soules potential using an experimental training data set that includes both local and global structural features from approximately 21 experimentally realized silica polymorphs, including high density phases and porous zeolites. We compare the performance of our machine-learned ML-Soules model with other high quality models including our recent machine-learned parametrization of BKS (ML-BKS), a machine-learned potential (GAP), as well as predictions of ab initio calculations with the highly fidelity SCAN functional. The ML-Soules accurately captures the relative energetic ordering of various polymorphs as well as their structural features at a significantly reduced computational expense. The ML-Soules model also reasonably captures the structure, density, and elastic constants of quartz, as well as metastable silica polymorphs. We further discuss the limitations of the Soules functional form and propose potential enhancements, including the incorporation of additional three-body terms and/or the utilization of different short-ranged functional forms to achieve greater accuracy for both global and local features in the modeling of silica while retaining low computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Convergent Protocols for Computing Protein–Ligand Interaction Energies Using Fragment-Based Quantum Chemistry

Fragment-based quantum chemistry methods offer a way to sidestep the steep nonlinear scaling of electronic structure calculations so that large molecular systems can be investigated using high-level methods. Here, we use fragmentation to compute protein–ligand interaction energies in systems with several thousand atoms, using a new software platform for managing fragment-based calculations that implements a screened many-body expansion. Convergence tests using a minimal-basis semiempirical method (HF-3c) indicate that two-body calculations, with single-residue fragments and simple hydrogen caps, are sufficient to reproduce interaction energies obtained using conventional supramolecular electronic structure calculations, to within 1 kcal/mol at about 1% of the computational cost. We also demonstrate that the HF-3c results are illustrative of trends obtained with density functional theory in basis sets up to augmented quadruple-ζ quality. Strategic deployment of fragmentation facilitates the use of converged biomolecular model systems alongside high-quality electronic structure methods and basis sets, bringing ab initio quantum chemistry to systems of hitherto unimaginable size. This will be useful for generation of high-quality training data for machine learning applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deciphering the Solvation Structure of Aqueous ZnCl 2 Solutions from X-ray Absorption Spectra Using the Interpretable Graph Neural Network

Machine learning (ML) provides powerful pathways for predicting spectroscopic observables from atomic structures, but its broader impact depends on making model predictions interpretable in terms of physical and chemical principles. Here, we introduce a physics-guided graph neural network (GNN) model that predicts Zn K-edge X-ray spectroscopy (XAS) spectra of aqueous ZnCl 2 solutions. Training data are generated from ab initio XAS calculations on molecular dynamics snapshots obtained using a machine learning interatomic potential. The GNN reproduces experimental spectra across concentrations from dilute (<0.1 m) to highly concentrated (30 m, “water-in-salt”) regimes and scales efficiently to large, disordered liquid systems beyond the reach of conventional ab initio approaches. Gradient-based attribution analysis reveals that the model learns physically meaningful structure-spectrum relationships. Ligand-specific attributions reflect orbital hybridization patterns and the origin of the excitations derived from the density functional theory. Bond-length attributions recover spectral shifts consistent with multiple-scattering theory. Finally, this work bridges data-driven prediction with electronic-structure theory, establishing a general paradigm for interpretable ML that links atomic structure, electronic structure, and spectroscopic observables.

25 ENERGY STORAGE↗

Predicting Partial Atomic Charges in Metal–Organic Frameworks: An Extension to Ionic MOFs

Molecular simulation is an invaluable tool to predict and understand the usage of metal–organic frameworks (MOFs) for gas storage and separation applications. Accurate partial atomic charges, commonly obtained from density functional theory (DFT) calculations, are often required to model the electrostatic interactions between the MOF and adsorbates, especially when the adsorbates have dipole or quadrupole moments, such as water and CO 2 . Machine learning (ML) models have been previously employed to predict partial charges and avoid the computational cost associated with DFT calculations. However, previous ML models suffer from small training data sets, which limit their scope of application. In this work, we introduce two novel machine learning models, PACMOF2-neutral and PACMOF2-ionic, aimed at predicting the density-derived electrostatic and chemical (DDEC6) partial atomic charges for both neutral and ionic MOFs. These models not only yield DFT-level accuracy at a fraction of the computational cost but also demonstrate a remarkable improvement in prediction of adsorption, as validated with grand canonical Monte Carlo simulations. Furthermore, the robustness and fast computational time of the PACMOF2 models, along with their transferability to other porous materials such as covalent organic frameworks and zeolites, underscores their potential in high-throughput screening of MOFs for diverse applications.

36 MATERIALS SCIENCE↗

Stable Simulation of the Community Atmosphere Model Using Machine‐Learning Physical Parameterization Trained With Experience Replay

In recent years, machine learning (ML) models have been used to improve physical parameterizations of general circulation models (GCMs). A significant challenge of integrating ML models into GCMs is the online instability when they are coupled for long‐term simulation. We present a new strategy that demonstrates robust online stability when the physical parameterization package of an atmospheric GCM is replaced by a deep ML model. The method uses experience replay with a multistep training scheme of the ML model in which the model's own output at the previous time step is used in the training. Predicted physics tendencies in the replay buffer with the most recent errors in the training iterations are reused, making the ML model learn from its own errors. The training method reduces the gap between the offline and online environments of the ML model. The method is used to train the ML model as the physical parameterization of the Community Atmosphere Model (CAM5) with training data from the Multi‐scale Modeling Framework high resolution simulations. Three 6‐year online simulations of the CAM5 are carried out by using the ML physics package. The simulated spatial distributions of precipitation, surface temperature and zonally averaged atmospheric fields demonstrate overall better accuracy than that of the standard CAM5 and benchmark model even without the use of additional physical constraints or tuning. This work is the first to demonstrate a solution to address the online instability problem in climate modeling with ML physics by using experience replay.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty quantification for neural network potential foundation models

Abstract For neural network potentials (NNPs) to gain widespread use, researchers must be able to trust model outputs. However, the blackbox nature of neural networks and their inherent stochasticity are often deterrents, especially for foundation models trained over broad swaths of chemical space. Uncertainty information provided at the time of prediction can help reduce aversion to NNPs. In this work, we detail two uncertainty quantification (UQ) methods. Readout ensembling, by finetuning the readout layers of an ensemble of foundation models, provides information about model uncertainty, while quantile regression, by replacing point predictions with distributional predictions, provides information about uncertainty within the underlying training data. We demonstrate our approach with the MACE-MP-0 model, applying UQ to the foundation model and a series of finetuned models. The uncertainties produced by the readout ensemble and quantile methods are demonstrated to be distinct measures by which the quality of the NNP output can be judged.

36 MATERIALS SCIENCE↗

Uncertainty quantification for misspecified machine learned interatomic potentials

The use of high-dimensional regression techniques from machine learning has significantly improved the quantitative accuracy of interatomic potentials. Atomic simulations can now plausibly target quantitative predictions in a variety of settings, which has brought renewed interest in robust means to quantify uncertainties. In many practical settings where model complexity is constrained (e.g., due to performance considerations), misspecification — the inability of any one choice of model parameters to exactly match all training data — is a key contributor to errors that is often disregarded. Here, we employ a recent misspecification-aware regression technique to quantify parameter uncertainties, which is then propagated to a broad range of phase and defect properties in tungsten. The propagation is performed through both brute-force resampling and implicit Taylor expansion. The propagated misspecification uncertainties robustly quantify and bound errors on a broad range of material properties. We demonstrate application to recent foundational machine learning interatomic potentials, accurately predicting and bounding errors in MACE-MPA-0 energy predictions across the diverse materials project database.

36 MATERIALS SCIENCE↗

Database and deep-learning scalability of anharmonic phonon properties by automated brute-force first-principles calculations

Understanding the anharmonic phonon properties of crystal compounds—such as phonon lifetimes and thermal conductivities—is essential for investigating and optimizing their thermal transport behaviors. These properties also impact optical, electronic, and magnetic characteristics through interactions between phonons and other quasiparticles and fields. In this study, we develop an automated first-principles workflow to calculate anharmonic phonon properties and build a comprehensive database encompassing more than 6500 inorganic compounds. Utilizing this dataset, we train a graph neural network model to predict thermal conductivity values and spectra from structural parameters, demonstrating a scaling law in which prediction accuracy improves with increasing training data size. High-throughput screening with the model enables the identification of materials exhibiting extreme thermal conductivities—both high and low. The resulting database offers valuable insights into the anharmonic behavior of phonons, thereby accelerating the design and development of advanced functional materials.

Ohnishi, Masato [University of Tokyo (Japan); Inst↗

The design space of E(3)-equivariant atom-centred interatomic potentials

Abstract Molecular dynamics simulation is an important tool in computational materials science and chemistry, and in the past decade it has been revolutionized by machine learning. This rapid progress in machine learning interatomic potentials has produced a number of new architectures in just the past few years. Particularly notable among these are the atomic cluster expansion, which unified many of the earlier ideas around atom-density-based descriptors, and Neural Equivariant Interatomic Potentials (NequIP), a message-passing neural network with equivariant features that exhibited state-of-the-art accuracy at the time. Here we construct a mathematical framework that unifies these models: atomic cluster expansion is extended and recast as one layer of a multi-layer architecture, while the linearized version of NequIP is understood as a particular sparsification of a much larger polynomial model. Our framework also provides a practical tool for systematically probing different choices in this unified design space. An ablation study of NequIP, via a set of experiments looking at in- and out-of-domain accuracy and smooth extrapolation very far from the training data, sheds some light on which design choices are critical to achieving high accuracy. A much-simplified version of NequIP, which we call BOTnet (for body-ordered tensor network), has an interpretable architecture and maintains its accuracy on benchmark datasets.

Computer Science↗

Efficiently predicting pressure-composition-temperature diagrams to discover low-stability metal hydrides

Quantitatively accurate computational predictions of metal hydride thermodynamics are challenging but critical for alloy performance optimization across a multitude of technological domains, including hydrogen storage, compression, purification, and getters. Recent machine learning approaches have demonstrated great success in this area, but can potentially suffer from several shortcomings since they rely on imbalanced experimental training data and can have poor out-of-distribution (ood) test performance. Here, in this study, we circumvent such pitfalls by developing a computationally efficient, first principles-based workflow for direct prediction of metal hydride phase equilibrium, i.e., the pressure-composition-temperature (PCT) diagram. We then demonstrate its utility on predicting low stability hydrides derived from compositionally complex C14 Laves phase AB2 alloys. Specifically, we computationally predict and then experimentally validate an AB 2 alloy series (z < 0.6 for Ti 2−z Zr z CrMnFeNi) with ideal hydriding thermodynamics for a two-stage metal hydride-based compressor for pressurizing boil off from liquefied hydrogen. Importantly, this study lays the groundwork for accurate and efficient discovery/optimization of ood, low-stability hydrides for which purely data-driven approaches lack sufficient accuracy.

08 HYDROGEN↗

Machine learning inversion of interatomic force constants from single-crystal inelastic neutron scattering

Atomic vibrations govern many macroscopic properties of materials, but experiments to comprehensively probe them remain challenging. Inelastic neutron scattering (INS) is a powerful technique to map phonon dispersions in crystals, especially when leveraging modern time-of-flight (ToF) spectrometers with large detectors. However, efficiently and robustly extracting interatomic force constants (FCs) parameterizing phonon dynamics from experimental spectra remains a bottleneck due to the complexity and high dimensionality of ToF INS datasets. Here, we present a machine learning approach for the direct inversion of FCs from single-crystal INS measurements. The framework leverages synthetic training data generated using universal machine-learned force fields and an efficient physics-based forward model. We benchmark two neural architectures–one emphasizing structured latent representation learning and the other direct, supervised spectral regression–across simulated datasets for two materials under idealized and noisy conditions. The latent-representation model is subsequently applied to experimental single-crystal INS data on germanium. The model is shown to reproduce FCs derived from both first-principles simulations and from iterative optimization, and furthermore achieves reliable inference even from sparse, single-orientation measurements representing short data acquisitions. Analysis of the learned latent space reveals semantically continuous and physically interpretable encodings that support strong cross-domain generalization. By bridging theoretical and experimental domains, we establish a path toward rapid inversion of experimental spectra and data-driven interpretation of temperature-dependent lattice dynamics.

42 ENGINEERING↗

Prototype-Wise Sensitivity Analysis of Urban Building Energy Simulation Surrogate Modeling Accuracy

Urban Building Energy Modeling (UBEM) is an important reference for urban energy-related policymaking. Because of the significant impact of urban microclimates on the energy simulation, UBEM requires simulations of many microclimate-prototype pairs. Surrogate modeling is commonly used to reduce the cost of simulation computations. In UBEM surrogate modeling, it is important to determine the percentage of microclimates related to a prototype used for generating surrogate model training data. This study analyzes the prototype-wise variations and sensitivities of surrogate model estimation accuracy to the microclimate sampling ratios. The results of the study can help determine the number of simulations used for generating surrogate modeling data, avoid redundant simulations, and reduce the computational cost for UBEM surrogate modeling and its time.

Pan, Xiyu↗

Shock Hugoniot calculations using on-the-fly machine learned force fields with ab initio accuracy

We present a framework for computing the shock Hugoniot using on-the-fly machine learned force field (MLFF) molecular dynamics simulations. In particular, we employ an MLFF model based on the kernel method and Bayesian linear regression to compute the free energy, atomic forces, and pressure, in conjunction with a linear regression model between the internal and free energies to compute the internal energy, with all training data generated from Kohn–Sham density functional theory (DFT). We verify the accuracy of the formalism by comparing the Hugoniot for carbon with recent Kohn–Sham DFT results in the literature. In so doing, we demonstrate that Kohn–Sham calculations for the Hugoniot can be accelerated by up to two orders of magnitude, while retaining ab initio accuracy. We apply this framework to calculate the Hugoniots of 14 materials in the FPEOS database, comprising 9 single elements and 5 compounds, between temperatures of 10 kK and 2 MK. We find good agreement with first principles results in the literature while providing tighter error bars. In addition, we confirm that the inter-element interaction in compounds decreases with temperature.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enabling accurate chemical modeling of shocked energetic materials using a machine learning interatomic potential

Understanding the complex chemistry of organic materials under dynamic compression is important for many applications, but it is challenging due to the large number of reactions occurring at various time scales. Here, in this study, we develop a machine learning potential based on Chebyshev polynomials to study the insensitive energetic material 1,3,5-triamino-2,4,6-trinitrobenzene (TATB) under detonation. We discuss a strategy for constructing diverse training data needed to capture the complex chemistry of TATB. Our potential demonstrates strong transferability across a wide range of thermodynamic conditions and other explosives, enabling accurate and reliable chemical modeling of organic materials under extreme conditions. The efficiency of our approach allows for simulations over several nanoseconds and for large system sizes, providing detailed insights into the chemistry of shocked TATB. The model accurately reproduces experimental Hugoniot equation of state data, and our simulations reveal the rapid formation of nitrogen-rich carbon clusters following shock. The methods and datasets developed here offer a robust framework for accurate chemical modeling of other shocked organic energetic materials.

Chemistry↗

Metal oxide candidates for thermochemical water splitting obtained with a generative diffusion model

Generative diffusion models (DMs) for inorganic crystalline materials are being actively investigated for their potential to expand the chemical and structural design spaces for known functional materials. Generative candidates are particularly useful for applications where few functional, let alone commercially viable, materials currently exist, such as metal oxides for thermochemical water-splitting, which have strict requirements for defect thermodynamics and host stability. Here, we critically examine generated metal oxides from the M ATTER G EN DM conditioned on select chemical systems for thermochemical water splitting applications. Perhaps most notably, we find that M ATTER G EN predicts a novel, thermodynamically stable, quinary metal oxide, Ba 2 SrInFeO 6 , although this compound represents an ordered and layered substitution within the same A 3 B 2 O 6 structural prototype as its two ternary end members. Detailed density functional theory calculations and spin configuration sampling for this material and its possible decomposition products—beyond what existed in M ATTER G EN training data—are required to quantitatively validate hull energy predictions and conclusions of stability. Furthermore, the material exhibits oxygen defect formation energies appropriate for thermochemical water splitting, warranting targeted investigation in an experimental validation campaign, along with other future M ATTER G EN candidates in this application space.

36 MATERIALS SCIENCE↗

Resimulation-based self-supervised learning for pretraining physics foundation models

Self-supervised learning (SSL) is at the core of training modern large machine learning models, providing a scheme for learning powerful representations that can be used in a variety of downstream tasks. However, SSL strategies must be adapted to the type of training data and downstream tasks required. We propose resimulation-based self-supervised representation learning (RS3L), a novel simulation-based SSL strategy that employs a method of resimulation to drive data augmentation for contrastive learning in the physical sciences, particularly, in fields that rely on stochastic simulators. By intervening in the middle of the simulation process and rerunning simulation components downstream of the intervention, we generate multiple realizations of an event, thus producing a set of augmentations covering all physics-driven variations available in the simulator. Using experiments from high-energy physics, we explore how this strategy may enable the development of a foundation model; we show how RS3L pretraining enables powerful performance in downstream tasks such as discrimination of a variety of objects and uncertainty mitigation. In addition to our results, we make the RS3L dataset publicly available for further studies on how to improve SSL strategies.

97 MATHEMATICS AND COMPUTING↗