Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning,”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

SRF cavity instability detection with machine learning at CEBAF

During the operation of the Continuous Electron Beam Accelerator Facility (CEBAF), one or more unstable superconducting radio-frequency (SRF) cavities often cause beam loss trips while the unstable cavities themselves do not necessarily trip off. The present RF controls for the legacy cavities report at only 1 Hz, which is too slow to detectfast transient instabilities during these trip events. These challenges make the identification of an unstable cavity out of the hundreds installed at CEBAF a difficult and time-consuming task. To tackle these issues, a fast data acquisition system (DAQ) for the legacy SRF cavities has been developed, which records the sample at 5 kHz. An unsupervised learning framework has been developed to identify anomalous SRF cavity behavior. We will discuss the present status of the DAQ system and our framework, along with recent successes in detecting anomalous cavity behavior. Overall, our method offers a practical solution for identifying unstable SRF cavities, contributing to increased beam availability and machine reliability.

Accelerator Physics↗

Using Active Learning to Rapidly Develop Machine Learned Diffusion Coefficients of CO 2 Conversion Reagents in Metal–Organic Frameworks

Here, we used a combined molecular dynamics/active learning (AL) approach to create machine learning models that can predict the diffusion coefficient of epichlorohydrin and chloropropene carbonate, the reactant and product of a common CO 2 cycloaddition reaction, in metal–organic frameworks (MOFs). Nanoporous MOFs are effective catalysts for the cycloaddition of CO 2 to epoxides. The diffusion rates within nanoporous catalysts can control the rate of reaction as the reactants and products must diffuse to the active sites within the MOF and then out of the nanoporous material for reusability. However, the diffusion process is routinely ignored when searching for new materials in catalytic applications. Here we verified improvement during the AL process by consistently tracking metrics on the same groups of MOFs to ensure consistency. Metal identity was found to have little impact on diffusion rates, while structural features like pore limiting diameter act as a threshold where a minimum value is needed for high diffusion rates. We identified the MOFs with the highest epichlorohydrin and chloropropene carbonate diffusion coefficients which can be used for further studies of reaction energetics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Physics-informed machine learning for building performance simulation-A review of a nascent field

Building performance simulation (BPS) is critical for understanding building dynamics and behavior, analyzing the performance of the built environment, optimizing energy efficiency, improving demand flexibility, and enhancing building resilience. However, conducting BPS is not trivial. Traditional BPS relies on accurate building energy models, which are primarily physics-based and heavily dependent on detailed building information, expert knowledge, and case-by-case model calibrations, significantly limiting their scalability. With the development of sensing technology and the increased availability of data, there is growing attention and interest in data-driven BPS. However, purely data-driven models often suffer from limited generalization ability and a lack of physical consistency, resulting in poor performance in real-world applications. To address these limitations, recent studies have begun integrating physics priors into data-driven models, a methodology known as physics-informed machine learning (PIML). PIML is an emerging field where its definitions, methodologies, evaluation criteria, application scenarios, and future directions remain open. To bridge those gaps, this study systematically reviews the state-of-the-art PIML for BPS, offering a comprehensive definition of PIML and comparing it to traditional BPS approaches regarding data requirements, modeling effort, performance, and computational cost. We also summarize the commonly used methodologies, validation approaches, application domains, available data sources, open-source packages, and testbeds. In addition, this study provides a general guideline for selecting appropriate PIML models based on BPS applications. Finally, this study identifies key challenges and outlines future research directions, providing a solid foundation and valuable insights to advance R&D of PIML in BPS.

Jiang, Zixin↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Hybrid Quantum Mechanical, Molecular Mechanical, and Machine Learning Potential for Computing Aqueous-Phase Adsorption Free Energies on Metal Surfaces

Performing reliable computer simulations of elementary processes occurring at metal–water interfaces is pivotal for novel catalyst design in sustainable energy applications. Computational catalyst design hinges on the ability to reliably and efficiently compute the potential energy surface (PES) of the system. Here, due to the large system sizes needed for studying processes at liquid water–metal interfaces, these systems can currently not be described using density functional theory (DFT). In this work, we used a hybrid quantum mechanical, molecular mechanical, and machine learning potential for studying the adsorption behavior of phenol, atomic hydrogen, 2-butanol, and 2-butanone on the (0001) facet of Ru under reducing conditions when Ru is not oxidized. Specifically, we describe the adsorbate and the surrounding metal atoms at the DFT level of theory. Here, we also considered the electrostatic field effect of the water molecules on adsorbate–metal interactions. Next, for the water–water and water–adsorbate interactions, we used established classical force fields. Finally, for the water–Ru surface interaction, for which no reliable force fields have been published, we used Behler–Parrinello high-dimensional neural network potentials (HDNNPs). Employing this setup, we used our explicit solvation for metal surface (eSMS) approach to compute the aqueous-phase effect on the low-coverage adsorption of selected molecules and atoms on the (0001) facet of Ru. In agreement with previous experimental and computational studies of oxygenated molecules over transition metal facets, we found that liquid water destabilizes the tested adsorbates on Ru(0001). Interestingly, our findings indicate that adsorbates on Ru are less affected by the presence of an aqueous phase than on other transition metals (e.g., Pt), highlighting the necessity of experimental investigations of Ru-based catalytic systems in liquid water.

Adsorption↗

A machine learning pipeline for identifying infiltration managed aquifer recharge locations from satellite imagery in the San Joaquin Valley, California

This study focuses on an agricultural region in California’s Central Valley, USA, where Managed Aquifer Recharge (MAR) is widely implemented to mitigate groundwater depletion under increasing water demand and climate variability. A deep learning and machine learning framework was developed to identify infiltration-MAR locations using satellite imagery and environmental data. The framework integrates surface water detection from Sentinel-2 imagery, geospatial delineation of water bodies, spatiotemporal tracking of water body dynamics, and supervised classification using meteorological, environmental, and topographic variables. The framework was applied to a 2379 km² study area southwest of Fresno, where 765 water bodies were detected, including 139 identified MAR sites based on publicly available datasets and expert knowledge. The classification model achieved an accuracy of 0.94 and an F1 score of 0.85. Feature importance analysis indicates that cropland, normalized difference vegetation index (NDVI), and evaporation are among the most influential predictors for infiltration-MAR. Notably, the framework suggests that engineered water management in infiltration-MAR systems can disrupt or even reverse the expected positive correlation between surface water extent and precipitation. These findings provide physically interpretable insights into the characteristics of existing infiltration-MAR facilities and demonstrate the potential of the proposed framework as a reproducible, interpretable, and potentially transferable tool for data-driven infiltration-MAR identification and inventory development under growing climatic and hydrological uncertainty.

Classification↗

Machine Learning Automation Pipeline

Machine Learning Automation Pipeline (MLAP) is a package to perform machine learning (ML) analysis in a step by step manner, starting with data extraction until analysis and prediction. The scripts provide the users option to chose an action such as "Extract", "Prep", and "Train" and numerous cases can be launched with just a single command. The inputs for each case are provided using a JSON file. The simulation results of several cases can be assessed using an automated process and analyzed for various metrics pertinent to ML analysis.

Jha, Pankaj↗

Systematic softening in universal machine learning interatomic potentials

Machine learning interatomic potentials (MLIPs) have introduced a new paradigm for atomic simulations. Recent advancements have led to universal MLIPs (uMLIPs) that are pre-trained on diverse datasets, providing opportunities for universal force fields and foundational machine learning models. However, their performance in extrapolating to out-of-distribution complex atomic environments remains unclear. In this study, we highlight a consistent potential energy surface (PES) softening effect in three uMLIPs: M3GNet, CHGNet, and MACE-MP-0, which is characterized by energy and force underprediction in atomic-modeling benchmarks including surfaces, defects, solid-solution energetics, ion migration barriers, phonon vibration modes, and general high-energy states. The PES softening behavior originates primarily from the systematically underpredicted PES curvature, which derives from the biased sampling of near-equilibrium atomic arrangements in uMLIP pre-training datasets. Our findings suggest that a considerable fraction of uMLIP errors are highly systematic, and can therefore be efficiently corrected. We argue for the importance of a comprehensive materials dataset with improved PES sampling for next-generation foundational MLIPs.

36 MATERIALS SCIENCE↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Characterizing Mesoscale Cellular Convection in Marine Cold Air Outbreaks With a Machine Learning Approach

Abstract During marine cold‐air outbreaks (MCAOs), when cold polar air moves over warmer ocean, a well‐recognized cloud pattern develops, with open or closed mesoscale cellular convection (MCC) at larger fetch over open water. The Cold‐Air Outbreaks in the Marine Boundary Layer Experiment provided a comprehensive set of ground‐based in situ and remote sensing observations of MCAOs at a coastal location in northern Norway. MCAO periods that unambiguously exhibit open or closed MCC are determined. Individual cells observed with a profiling Ka‐band radar are identified using a watershed segmentation method. Using self‐organizing maps (SOMs), these cells are then objectively classified based on the variability in their vertical structure. The SOM nodes contain some information about the location of the cell transect relative to the center of the MCC. This adds classification noise, requiring numerous cell transects to isolate cell dynamical information. The SOM‐based classification shows that comparatively intense convection occurs only in open MCC. This convection undergoes an apparent lifecycle. Developing cells are associated with stronger updrafts, large spectrum width, larger amounts of liquid water, lower surface precipitation rates, and lower cloud tops than mature and weakening cells. The weakening of these cells is associated with the development of precipitation‐induced cold pools. The SOM classification also reveals less intense convection, with a similar lifecycle. More stratiform vertical cloud structures with weak vertical motions are common during closed MCC periods and are separated into precipitating and non‐precipitating stratiform cores. Convection is observed only occasionally in the closed MCC environment.

Meteorology & Atmospheric Sciences↗

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

Learning to Branch with Interpretable Machine Learning Models

This presentation describes an algorithm for applying machine learning to branching to speed up the solution of integer optimization problems. These problems are challenging and solved multiple times a day by power systems operators. We show that our approach speeds up a widely used open-source optimization solver.

Bayramoglu, Selin↗

When more data hurts: Optimizing data coverage while mitigating diversity-induced underfitting in an ultrafast machine-learned potential

Machine-learned interatomic potentials (MLIPs) are becoming an essential tool in materials modeling. However, optimizing the generation of training data used to parametrize the MLIPs remains a significant challenge. This is because MLIPs can fail when encountering local environments too different from those present in the training data. The difficulty of determining a priori the environments that will be encountered during molecular dynamics simulation necessitates diverse, high-quality training data. Here, this study investigates how training data diversity affects the performance of MLIPs using the Ultra-Fast force field (UF 3 ) to model amorphous silicon nitride. We employ expert and autonomously generated data to create the training data and fit four force field variants to subsets of the data. Our findings reveal a critical balance in training data diversity: insufficient diversity hinders generalization, while excessive diversity can exceed the MLIP's learning capacity, reducing simulation accuracy. Specifically, we found that the UF 3 variant trained on a subset of the training data, in which nitrogen-rich structures were removed, offered vastly better prediction and simulation accuracy than any other variant. By comparing these UF 3 variants, we highlight the nuanced requirements for creating accurate MLIPs, emphasizing the importance of application-specific training data to achieve optimal performance in modeling complex material behaviors.

ab initio molecular dynamics↗

Impact of classical statistics on thermal conductivity predictions of BAs and diamond using machine learning molecular dynamics

Machine learning interatomic potentials (MLIPs) have greatly enhanced molecular dynamics (MD) simulations, achieving near-first-principles accuracy in thermal conductivity studies. In this work, we reveal that this accuracy, observed in BAs and diamond at sub-Debye temperatures, stems from an accidental error cancelation: classical statistics overestimates specific heat while underestimating phonon lifetimes, balancing out in thermal conductivity predictions. However, this balance is disrupted when isotopes are introduced, leading MLIP-based MD to significantly underpredict thermal conductivity compared to experiments and quantum statistics-based Boltzmann transport equation. This discrepancy arises not from classical statistics affecting phonon–isotope scattering rates but from its impact on the interplay between phonon–isotope and phonon–phonon scattering in the normal scattering-dominated BAs and diamond. In conclusion, this work underscores the limitations of MLIP-based MD for thermal conductivity studies at sub-Debye temperatures.

36 MATERIALS SCIENCE↗

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Characterizing Seasonal Variation of the Atmospheric Mixing Layer Height Using Machine Learning Approaches

As machine learning becomes more integrated into atmospheric science, XGBoost has gained popularity for its ability to assess the relative contributions of influencing factors in the atmospheric boundary layer height. To examine how these factors vary across seasons, a seasonal analysis is necessary. However, dividing data by season reduces the sample size, which can affect result reliability and complicate factor comparisons. To address these challenges, this study replaces default parameters with grid search optimization and incorporates cross-validation to mitigate dataset limitations. Using XGBoost with four years of data from the atmospheric radiation measurement (ARM) (Southern Great Plains (SGP) C1 site, cross-validation stabilizes correlation coefficient fluctuations from 0.3 to within 0.1. With optimized parameters, the R value can reach 0.81. Analysis of the C1 site reveals that the relative importance of different factors changes across seasons. Lower tropospheric stability (LTS, ~0.53) is the dominant factor at C1 throughout the year. However, during DJF, latent heat flux (LHF, 0.44) surpasses LTS (0.22). In SON, LTS (0.58) becomes more influential than LHF (0.18). Further comparisons among the four long-term SGP sites (C1, E32, E37, and E39) show seasonal variations in relative importance. Notably, during JJA, the differences in the relative importance of the three factors across all sites are lower than in other seasons. This suggests that boundary layer development in the summer is not dominated by a single factor, reflecting a more intricate process likely influenced by seasonal conditions such as enhanced convective activity, higher temperatures, and humidity, which collectively contribute to a balanced distribution of parameter impacts. Furthermore, the relative importance of LTS gradually increases from morning to noon, indicating that LTS becomes more significant as the boundary layer approaches its maximum height. Consequently, the LTS in the early morning in autumn exhibits greater relative importance compared to other seasons. This reflects a faster development of the mixing layer height (MLH) in autumn, suggesting that it is easier to retrieve the MLH from the previous day during this period. The findings enhance understanding of boundary layer evolution and contribute to improved boundary layer parameterization.

54 ENVIRONMENTAL SCIENCES↗