Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning (ML)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Rapid Coal-Ash Characterization using Geophysical Methods & Machine Learning

Coal combustion products (CCP) are challenging to delineate in heterogeneous field settings. Conventional methods (test pits, coring, and laboratory analyses) are labor-intensive, slow, invasive, and provide sparse spatial coverage. This study evaluates whether rapid non-invasive geophysical screening methods—induced polarization (IP), magnetic susceptibility, and nuclear magnetic resonance (NMR) —combined with surface colorimetry (RGB_24), can discriminate CCP-soil mixtures and provide reliable estimates of CCP content. Laboratory measurements were collected on five CCP-soil mixtures (series) and modeled using (i) a linear baseline, (ii) a calibrated non-linear (power-mean) model, and (iii) a machine-learning (ML) Random Forest approach, with validation via leave-one-series-out and site-specific tests. Across the five series, individual signals—particularly IP and magnetic susceptibility—were strongly predictive of ash content but were consistently outperformed by combined models. The pooled calibrated non-linear and ML models captured the observed non-linearity and achieved high accuracy and precision, improving on linear fits. Colorimetry showed the weakest direct relationship with ash content for the tested samples but improved performance when included in multi-signal models. At pre-selected 3.5% decision threshold, calibrated and ML approaches yielded near-perfect classification (Matthews correlation coefficient ˜ 1), suggesting strong practical operability for field screening. Additionally, field-analog tests highlighted the role of endmembers—accuracy declined without access to end-member measurements but was largely recovered by collecting a minimal labeled pair for local recalibration. With end members, accuracy remained high. Globally trained models performed well on three operational unknowns; however, series-specific refits provided the most accurate predictions. Overall, these results highlight the potential of combining rapid geophysics and minimal local calibration for improved coal-ash delineation.

Peshtani, Klaudio↗

Enhancing 2D hydrodynamic flood models through machine learning and urban drainage integration

Two-dimensional hydrodynamic flood models are commonly employed for simulating flood extent and inundation depth. However, the influence of urban drainage network (UDN) is frequently overlooked in these models, potentially compromising their accuracy. Furthermore, the expensive computational costs and longer processing times make them challenging for large-scale hydrodynamic simulation. To address these challenges, this paper develops a machine learning (ML)-driven emulator for an open-source flood model, the Two-dimensional Runoff Inundation Toolkit for Operational Needs (TRITON). A TRITON-ML Emulator (TR-Emulator) that utilizes Convolutional Long Short-Term Memory is developed to capture the spatiotemporal features of flood events based on the outputs from TRITON. We further enhance the emulator by integrating UDN parameters (TR-UDN), such as the flow capacity of drainage pipes, pipe size, and pipe length, via an ML stacking technique to improve the water surface elevation (WSE) simulation. Hurricane Harvey 2017 in Houston, TX is used as the case study. We compare WSE results from TRITON, TR-Emulator, TR-UDN, and the United States Geological Survey (USGS) observations to evaluate the performance of these models. The results indicate that the TR-Emulator effectively replicates the WSE simulated by TRITON. Additionally, TR-UDN performs well in capturing WSE patterns and peak flows, aligning more closely with USGS observations, except in areas with milder slopes where conveyance discrepancies are observed. We further test the generalizability of our ML-based models using another smaller event. This paper shows that the TR-Emulator is effective for users and engineers to emulate a 2D hydrodynamic model, and the enhanced version of the TR-Emulator, TR-UDN, can be an efficient tool for predicting WSEs during urban flooding.

54 ENVIRONMENTAL SCIENCES↗

Low Activity Waste Glass Optimization with Property Models from Machine Learning, Part 2: Experimental Validation and Active Learning

The United States Department of Energy is responsible for managing legacy nuclear waste stored in underground tanks at the Hanford Site. To treat the waste, it is planned as the current baseline to separately vitrify low-activity waste (LAW) and high-level waste fractions. Previously, machine learning (ML) based glass property models (e.g., chemical durability, viscosity, electrical conductivity and SO3 solubility) were developed with prediction uncertainties. A waste glass optimization approach was then established to enable the capability of using these ML models in LAW glass formulation. In this study, the previous ML models were first experimentally validated, and the results were incorporated back into the database to update the ML models. The updated models and formulations showed increased waste loading while reducing the failure rate, demonstrating improved predictive accuracy, reduced uncertainties, and the effectiveness of active learning in guiding high-dimensional, nonlinear LAW glass design. This represents the first experimental validation of ML based LAW glass formulation, with practical benefits such as higher waste loading, shorter mission duration, and lower operational risk.

Lu, Xiaonan (ORCID:0000000179708148)↗

MOOSE ProbML: Parallelized probabilistic machine learning and uncertainty quantification for computational energy applications

Here, this paper presents the development and demonstration of massively parallel probabilistic machine learning (ML) and uncertainty quantification (UQ) capabilities within the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source computational platform for parallel finite element and finite volume analyses. In addressing the computational expense and uncertainties inherent in complex multiphysics simulations, this paper integrates Gaussian process (GP) variants, active learning, Bayesian inverse UQ, adaptive forward UQ, Bayesian optimization, evolutionary optimization, and Markov chain Monte Carlo (MCMC) within MOOSE. It also elaborates on the interaction among key MOOSE systems — Sampler, MultiApp, Reporter, and Surrogate — in enabling these capabilities. The modularity offered by these systems enables development of a multitude of probabilistic ML and UQ algorithms in MOOSE. Example code demonstrations include parallel active learning and parallel Bayesian inference via active learning. The impact of these developments is illustrated through five applications relevant to computational energy applications: UQ of nuclear fuel fission product release, using parallel active learning Bayesian inference; very rare events analysis in nuclear microreactors using active learning; advanced manufacturing process modeling using multi-output GPs (MOGPs) and dimensionality reduction; fluid flow using deep GPs (DGPs); and tritium transport model parameter optimization for fusion energy, using batch Bayesian optimization. These capabilities are part of the MOOSE framework.

97 - MATHEMATICS AND COMPUTING↗

Dense autoencoders, clustering techniques, and semi-supervised learning for HPGe $γ$-spectra

Classifying high-resolution gamma spectra by their isotopic content is an essential task in nuclear forensics and other applications. Traditional analysis methods are often time-intensive, but machine learning (ML) may help analysts quickly process many spectra. Such methods tend to rely on abundant, well-labeled data for training. Historical gamma data exists in various fields but is not uniformly useful for supervised ML due to inconsistent labeling. Here, to address some of these challenges, we present a method to classify and organize unlabeled data from high-purity germanium detectors using an autoencoding neural network (autoencoder). We trained dense autoencoders to compress gamma data into latent representations that enable efficient data characterization. By clustering the encoded spectra or lower-dimensional mappings of them, we identified and removed portions of over-abundant data categories, resulting in a more balanced dataset and improved autoencoder performance. This encoding and clustering pipeline also enabled the organization of spectra into self-consistent categories. Finally, we found that encoded representations showed potential as inputs for semi-supervised learning of nuclide identification (NID) labels, achieving an average F1 score of 0.85 ± 0.03 when mapping encodings to a set of 65 isotope labels.

Autoencoders↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

Designing resilient IoT and Edge Computing with federated tinyML

The rapid growth of the Internet of Things (IoT) and Edge Computing (EC) has brought significant conveniences to modern society but has also greatly expanded the cyber attack surfaces, particularly as these technologies are being increasingly integrated into critical systems such as power grids, healthcare, and smart homes. Here, to improve IoT/EC’s cybersecurity posture, we leveraged Artificial Intelligence (AI) and Machine Learning (ML) by employing tinyML to monitor voluminous IoT data for cyber threats while addressing devices’ resource constraints, and utilizing Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. Building on our three-layer architecture combining tinyML and FL to enhance autonomous cyber attack detection, this paper demonstrated that the architecture improves detection accuracy, reduces resource consumption, and enables lightweight, secure IoT device monitoring. These results were validated using the public N-BaIoT dataset as well as real IoT network traffic data collected under multiple attack scenarios from our testbeds. Additionally, we introduced an enhanced FL methodology with a novel preprocessing stage, including federated feature selection and global preprocessor construction, to address IoT/EC data heterogeneity. We developed a physical IoT testbed for attack simulations and data collection, implemented a tinyML-powered detector for realistic model validation, and also built a virtual testbed for scalable evaluations of FL models across diverse network environments.

Cognitive cyber↗

On the effectiveness of neural operators at zero-shot weather downscaling

Machine-learning (ML) methods have shown great potential for weather downscaling. These data-driven approaches provide a more efficient alternative for producing high-resolution weather datasets and forecasts compared to physics-based numerical simulations. Neural operators, which learn solution operators for a family of partial differential equations, have shown great success in scientific ML applications involving physics-driven datasets. Neural operators are grid-resolution-invariant and are often evaluated on higher grid resolutions than they are trained on, i.e., zero-shot super-resolution. Given their promising zero-shot super-resolution performance on dynamical systems emulation, we present a critical investigation of their zero-shot weather downscaling capabilities, which is when models are tasked with producing high-resolution outputs using higher upsampling factors than are seen during training. To this end, we create two realistic downscaling experiments with challenging upsampling factors (e.g., 8x and 15x) across data from different simulations: the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) and the Wind Integration National Dataset Toolkit. While neural operator-based downscaling models perform better than interpolation and a simple convolutional baseline, we show the surprising performance of an approach that combines a powerful transformer-based model with parameter-free interpolation at zero-shot weather downscaling. We find that this Swin-Transformer-based approach mostly outperforms models with neural operator layers in terms of average error metrics, whereas an Enhanced Super-Resolution Generative Adversarial Network-based approach is better than most models in terms of capturing the physics of the ground truth data. We suggest their use in future work as strong baselines.

17 WIND ENERGY↗

Machine Learning-Guided Optimization of SABRE Hyperpolarization for α-Ketoglutarate in Acetone–Water

Signal amplification by reversible exchange (SABRE) is a hyperpolarization method that polarizes target nuclei of metabolites quickly and efficiently. Recent SABRE advances, including Ace-SABRE, yield biocompatible, aqueous solutions of hyperpolarized markers for metabolic monitoring. Building on recent advancements, expanding the substrate scope of Ace-SABRE is desirable. However, SABRE polarization is sensitive to many different parameters; therefore, traditional optimization approaches are experimentally time-consuming. In this proof-of-concept application of machine learning (ML), Bayesian optimization (BO) is used for four important input parameters to model the complex SABRE dynamics while saving experimental time. The presented ML model also provides chemical insights that enable predictions of sample compositions for increased polarization levels. In this paper, we transition from an original average free polarization of p = ∼0.90% to a maximum observed free polarization of p = ∼6.6% for 1- 13 C alpha-ketoglutarate (AKG) with 13 C at natural abundance, utilizing both direct outputs as well as chemical insights revealed by the ML model.

Catalysts↗

Understanding the Effect of Sample Geometry on Temperature Distribution during Optical Floating Zone Crystal Growth in Vacuum Environment through Heat Transfer Modeling

Optical floating zone furnaces (OFZ) have had a transformative impact on fundamental science due to their ability to rapidly produce large single crystals of a wide variety of complex materials. However, a quantitative understanding of the OFZ growth environment is generally lacking due to the difficulty of measuring the local sample temperatures during OFZ growth, as well as to the general lack of information about the temperature-dependent physical parameters needed to model heat transfer. To overcome these challenges, we apply a physics-based heat transfer model, parametrized by measurements from synchrotron experiments and a machine-learning (ML) algorithm, to simulate the temperature distributions of samples heated in an OFZ furnace in a vacuum environment. This model is used to quantitatively understand how the sample maximum temperature and temperature gradient (key parameters that influence the success of crystal growth) are affected by the rod size, rod shape, and heat-zone position on the rod. The results of this study can be applied to make informed decisions on how crystal growth parameters can be tuned to modify temperature profiles and to optimize crystal growth outcomes even when data on internal sample temperature profiles (e.g., those obtained through in situ synchrotron experiments) are not accessible.

36 MATERIALS SCIENCE↗

Accelerating the Discovery of New, Single Phase High Entropy Ceramics via Active Learning

High-entropy ceramics have garnered interest due to their remarkable hardness, compressive strength, thermal stability, and fracture toughness; yet the discovery of new high-entropy ceramics (out of a tremendous number of possible elemental permutations) still largely requires costly, inefficient, trial-and-error experimental and computational approaches. The entropy forming ability (EFA) factor was recently proposed as a computational descriptor that positively correlates with the likelihood that a 5-metal high-entropy carbide (HECs) will form the desired single phase, homogeneous solid solution; however, discovery of new compositions is computationally expensive. If you consider 8 candidate metals, the HEC EFA approach uses 49 optimizations for each of the 56 unique 5-metal carbides, requiring a total of 2744 costly density functional theory calculations. Here, we describe an orders-of-magnitude more efficient active learning (AL) approach for identifying novel HECs. To begin, we compared numerous methods for generating composition-based feature vectors (e.g., magpie and mat2vec), deployed an ensemble of machine learning (ML) models to generate an average and distribution of predictions, and then utilized the distribution as an uncertainty. Here we then deployed an AL approach to extract new training data points where the ensemble of ML models predicted a high EFA value or was uncertain of the prediction. Our approach has the combined benefit of decreasing the amount of training data required to reach acceptable prediction qualities and biases the predictions toward identifying HECs with the desired high EFA values, which are tentatively correlated with the formation of single phase HECs. Using this approach, we increased the number of 5-metal carbides screened from 56 to 15,504, revealing 4 compositions with record-high EFA values that were previously unreported in the literature. Our AL framework is also generalizable and could be modified to rationally predict optimized candidate materials/combinations with a wide range of desired properties (e.g., mechanical stability, thermal conductivity).

36 MATERIALS SCIENCE↗

Evaluating Material Design Principles for Calcium-Ion Mobility in Intercalation Cathodes

Multivalent-ion batteries offer an alternative to Li-based technologies, with the potential for greater sustainability, improved safety, and higher energy density, primarily due to their rechargeable system featuring a passivating metal anode. Although a system based on the Ca 2+ /Ca couple is particularly attractive given the low electrochemical plating potential of Ca 2+ , the remaining challenge for a viable rechargeable Ca battery is to identify Ca cathodes with fast ion transport. In this work, a high-throughput computational pipeline is adapted to (1) discover novel Ca cathodes in a largely unexplored space of empty intercalation hosts and (2) develop material design rules for Ca-ion mobility. One candidate from the screening, W 2 O 3 (PO 4 ) 2 , is confirmed to have a low Nudged Elastic Band (NEB) barrier of 168 meV within a one-dimensional (1D) ion percolation topology. This candidate is subsequently synthesized and electrochemically tested, achieving reversible Ca cycling with a capacity of 25 mA h/g. To further accelerate the screening for promising Ca intercalation electrodes, machine learning (ML) Random Forest (RF) and Extreme Gradient Boosting (XGB) classification models are created with local environment descriptors based on a large, structurally and chemically diverse dataset of minimum energy pathways, spanning over 5,000 density functional theory (DFT) site energy calculations. Accuracies of 92% are achieved, material design metrics are quantified, ML force-fields are leveraged in an accelerated iteration of the screening, and a total of 27 novel Ca cathode materials are highlighted for further investigation.

25 ENERGY STORAGE↗

Data-Driven Discovery and Experimental Validation of Solvent Polarity Effects on Conjugated Polymer Solution-to-Film Assembly Pathways

Understanding how solvent properties influence the solution-to-film assembly of conjugated polymers remains a critical challenge due to the complex and intertwined nature of polymer–solvent interactions. In this study, we integrate a data-driven framework with experimental validation to identify key parameters influencing the assembly and performance of poly[2,5-(2-octyldodecyl)-3,6-diketopyrrolopyrrole-alt-5,5-(2,5-di(thien-2-yl)thieno[3,2-b]thiophene)] (DPP-DTT) in organic field-effect transistors (OFETs). A machine learning (ML) approach identified the normalized Reichardt polarity parameter (E T N ) as a significant descriptor correlated with DPP-DTT hole mobility (μ). Systematic DPP-DTT devices fabricated using solvents across a wide E T N range revealed that higher E T N solvents yield enhanced μ. To elucidate the structural origins of high μ, we conducted comprehensive analyses using UV–vis–NIR spectroscopy and grazing incidence wide angle X-ray scattering (GIWAXS) measurements. The results revealed that films processed from high E T N solvents exhibit reduced paracrystallinity. By analyzing the solution-state behavior using optical microscopy and solution WAXS, we revealed polymer solubility differences in the various solvents and associated distinct polymer assembly pathways, elucidating why the high E T N solvent produces long-range ordered films. Notably, the high E T N solvent shows a pronounced preference for liquid-crystal (LC)-mediated assembly, providing a mechanistic explanation for the enhanced structural order. Therefore, these results demonstrate that solvent polarity, as evaluated by E T N , serves as an important parameter that plays a significant role in the DPP-DTT assembly pathway and resultant solid-state morphology. This work provides a strategy for integrating data science with experiments to identify critical parameters associated with complex polymer systems and helps guide rational process design for high-performance organic electronics.

36 MATERIALS SCIENCE↗

Machine learning-based bias-corrected future projections of ozone concentrations from a chemistry-climate model

Reliable projection of future near-surface ozone is crucial for air quality management and health risk assessment. However, potential biases in spatial distribution, magnitude and trends in ozone concentrations simulated by global chemistry-climate models limit their applicability in regional-scale evaluations. In this study, LightGBM, a machine learning (ML) algorithm is applied to correct biases in CESM2-simulated ozone concentrations over China, the United States and Europe and calibrate future ozone projections under two diverse Shared Socioeconomic Pathways (SSP1-2.6 and SSP5-8.5) scenarios from 2020 to 2060. The ML-based correction significantly improves the spatial distribution and reduces the model bias by 40%–60%. It also reverses the potentially incorrect trend of ozone change under SSP1-2.6 in eastern China. When applying ML-based bias correction to CESM2 future projections, warm season mean ozone concentrations decrease across China, the United States, and Europe by –13.5, –17.9, and –13.7 µg/m³, respectively, between 2020 and 2060 in SSP1-2.6, while they increase by 9.4, 2.0, and 5.2 µg/m³ in SSP5-8.5. Decomposition analysis show that changes in anthropogenic emissions dominate future ozone changes in both scenarios, while strong climate penalty from ozone changes occurs in polluted eastern China and climate benefit is found in western China, the United States and Europe under SSP5-8.5. These findings demonstrate the value of combining ML with chemistry-climate models to produce more accurate air quality projections, thereby informing more effective and region-specific environmental protection strategies.

Chemistry Model↗

Machine Learning-Driven Solvent Screening for Biobased 2,3-Butanediol Extraction

Biobased 2,3-butanediol (2,3-BDO) is a valuable biomass-derived chemical due to its versatility in being transformed into a wide variety of products. However, the separation and purification of 2,3-BDO from fermentation broth remain a significant challenge owing to its high boiling point and hydrophilic nature. Herein, we developed a machine learning (ML)-based screening workflow that uses molecular calculations as training data and requires only a small number of experimental measurements for validation to identify alternative solvent candidates for the liquid–liquid extraction (LLE) of 2,3-BDO from aqueous solution. In particular, 130 density functional theory (DFT) calculations with the implicit solvation method not only built a correlation between the computational partition coefficient and the experimental distribution coefficient of 2,3-BDO but also parameterized an Extra-Trees ML model to screen the distribution coefficient for a wider range of 6717 organic solvents. The experimental measurements of only 24 solvents were needed to validate the computational results. A list of 50 prioritized solvents was proposed for 2,3-BDO LLE, and seven additional experimental measurements were conducted to further verify our selected solvents. The impact of the extraction temperature and solvent-to-feed ratio was also investigated for selected solvents in experiments. Furthermore, this work suggested alternative solvents for 2,3-BDO LLE and proposed a versatile workflow that requires fewer experiments and can be applied to a broader range of LLE studies.

Extraction↗

Search for Stable and Low-Energy Ce–Co–Cu Ternary Compounds Using Machine Learning

Cerium-based intermetallics have garnered significant research attention as potential new permanent magnets. In this study, we explore the compositional and structural landscape of Ce−Co−Cu ternary compounds using a machine learning (ML)- guided framework integrated with first-principles calculations. We employ a crystal graph convolutional neural network (CGCNN), which enables efficient screening for promising candidates, significantly accelerating the material discovery process. With this approach, we predict five stable compounds, Ce 3 Co 3 Cu, CeCoCu 2 , Ce 12 Co 7 Cu, Ce 11 Co 9 Cu, and Ce 10 Co 11 Cu 4 , with formation energies below the convex hull, along with hundreds of low-energy (possibly metastable) Ce−Co−Cu ternary compounds. Firstprinciples calculations reveal that several structures are both energetically and dynamically stable. Notably, two Co-rich low-energy compounds, Ce 4 Co 33 Cu and Ce 4 Co 31 Cu 3 , are predicted to have high magnetizations.

Chemical structure↗

A Comprehensive Machine Learning Model for Metal–Ligand Binding Prediction: Applications in Chemistry and Biology

A machine-learning (ML) model that predicts metal–ligand binding constants was developed using the open-source Chemprop software. The model was trained on over 30,000 experimental log K 1 values, which include both protonation and metal–ligand stability constants, comprising over 3500 ligands and 10 2 metal ions from 73 total elements, thus generalizing beyond existing limited approaches, which focus only on specific metals or ligand families. The best-performing model included a combination of SMILES-based molecular representations along with descriptors for the metal ion and experimental conditions. It had an external test R 2 value of 0.942, and MAE value of 0.834. A “SMILES-only” simpler version also produced accurate predictions and preserved the binding trends, serving as a quick and easily accessible alternative for users without computational expertise. The SMILES-only model performed comparably to density functional theory (DFT) calculations but utilized a fraction of the computational resources. The model was successfully applied across diverse domains, including bioinorganic chemistry, heavy metal remediation, and sensor development and demonstrated its effectiveness as a rapid and reliable screening tool for both academic and industrial uses.

Ligands↗

Similarity Metric for Data Optimization and Efficient Training of Reactive Machine Learning Force Fields for Hydrocarbon Radiolysis

Radiolysis is a common approach to sterilize polymers, chemically modify them for upcycling, and accelerate their decomposition for recycling purposes. Reactive molecular dynamics (MD) simulations provide a powerful tool to generate atomic-level trajectories of the reactive processes and quantify radiolytic chemical degradation pathways. For this, machine learning (ML) surrogate models for reactive force fields with quantum mechanical accuracy are now widely used, which require ML training data sets that can provide information on atomic environments for target chemical systems. However, radiolysis chemistry can be highly complex and diverse, which poses significant challenges for generating training data to parametrize ML models. In this regard, we developed a method for optimizing the training data set using a cosine similarity metric to help guide training set selection for radiolysis of polyethylene, a model hydrocarbon polymer, as well as to enhance the transferability of our reactive ML force field (MLFF) to a variety of molecular and polymeric systems. Our approach performs atom-by-atom comparisons between local atomic environments to pinpoint important data points associated with rare and localized events, such as radiolysis damage within structures. We apply this approach to train the Chebyshev Interaction Model for Efficient Simulation (ChIMES) MLFF model, which expresses the atomic interaction potentials in terms of linear combinations of many-body Chebyshev polynomials. We first show that our method can reduce our training set size by ∼70% while improving overall accuracy compared to more standard MD model fitting approaches. We then validate our optimum model against diverse hydrocarbon simulation data, including simple alkanes and systems with unsaturated carbon bonds, over a wide range of thermodynamic conditions. Finally, we use our ChIMES model to perform MD simulations of radiolytic damage with large-scale systems that help avoid system size effects. Overall, our approach yields an MD force field that retains most of the accuracy of the underlying quantum method while yielding many orders of improvement in computational efficiency. In conclusion, our efforts will have impact on future hydrocarbon polymer radiolysis studies, where the chemical details of the polymer–radiation interactions can have a strong effect on the resulting products observed in experiments.

Hydrocarbons↗