Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning (ML)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Insights into Prismatic Loop Formation in Irradiated Fe–Cr Alloys from Hypothesis-Driven Active Learning and Causal Analysis

Neutron and electron irradiation experimental studies conducted on body-centered cubic Fe and Fe–Cr alloys have established two prismatic dislocation loop populations, which have Burgers vectors of either a/2$\langle$111$\rangle$ or a$\langle$100$\rangle$. Here, the loop formation depends on factors such as dose (D), dose rate (D rt ), temperature (T), chromium content (Cr%), and other alloying elements. Hence, it is important to understand how irradiation-induced dislocation loops evolve conditional upon the loop characteristics, such as loop density (DD), average loop size d̅, and irradiation parameters (D, D rt , T, and irradiation type), which is still an active area of research. To understand these complex structure–property relationships, machine learning (ML) is employed in a three-step approach. This includes imputing missing data with a k-nearest neighbor, generating functionalized features, and assessing feature importance with random forest classification and regression. Physics-based features are incorporated in a hypothesis-driven active learning scheme to overcome data unavailability challenges. Insights obtained from ML models (i) to categorize dislocation loop types, show the highest correlation with d̅; (ii) Log(DD), obtained through mathematical formulations involving D, Cr%, d̅, and T (e.g., Log(DD) ~ D + exp(-Cr%) + 1/d̅ and log(DD) ~ D + exp(-Cr%) + 1/T). Hypothesis-driven active learning is able to predict Log(DD) in which the experimental date is not known. Causal models verify cause–effect relationships for dislocation loop classification and irradiation factors in FeCr alloys.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Active and Transfer Learning of High-Dimensional Neural Network Potentials for Transition Metals

Classical molecular dynamics (MD) simulations represent a very popular and powerful tool for materials modeling and design. The predictive power of MD hinges on the ability of the interatomic potential to capture the underlying physics and chemistry. There have been decades of seminal work on developing interatomic potentials, albeit with a focus predominantly on capturing the properties of bulk materials. Such physics-based models, while extensively deployed for predicting the dynamics and properties of nanoscale systems over the past two decades, tend to perform poorly in predicting nanoscale potential energy surfaces (PESs) when compared to high-fidelity first-principles calculations. These limitations stem from the lack of flexibility in such models, which rely on a predefined functional form. Machine learning (ML) models and approaches have emerged as a viable alternative to capture the diverse size-dependent cluster geometries, nanoscale dynamics, and the complex nanoscale PESs, without sacrificing the bulk properties. Here, in this study, we introduce an ML workflow that combines transfer and active learning strategies to develop high-dimensional neural networks (NNs) for capturing the cluster and bulk properties for several different transition metals with applications in catalysis, microelectronics, and energy storage, to name a few. Our NN first learns the bulk PES from the high-quality physics-based models in literature and subsequently augments this learning via retraining with a higher-fidelity first-principles training data set to concurrently capture both the nanoscale and bulk PES. Our workflow departs from status-quo in its ability to learn from a sparsely sampled data set that nonetheless covers a diverse range of cluster configurations from near-equilibrium to highly nonequilibrium as well as learning strategies that iteratively improve the fingerprinting depending on model fidelity. All the developed models are rigorously tested against an extensive first-principles data set of energies and forces of cluster configurations as well as several properties of bulk configurations for 10 different transition metals. Our approach is material agnostic and provides a methodology to transfer and build upon the learnings from decades of seminal work in molecular simulations on to a new generation of ML-trained potentials to accelerate materials discovery and design.

36 MATERIALS SCIENCE↗

..delta..-Learning of High-Fidelity Electronic Structure Using Graph Neural Networks with Modified Node-Level Features

In this work, we present a ..delta..-learning approach for predicting the eigenvalues calculated with the hybrid functional HSE06 (..epsilon..nkHSE) for a set of metal and nitrogen doped graphene catalysts (MNCs) from Perdew-Burke-Ernzerhof (PBE) inputs. The model presented here incorporates electronic scalar features along with structural information in a graph neural network (GNN). In particular, the PBE eigenvalues for different bands and k-points and orbital-resolved projectors are combined with the applied potential as node-level features along with structural information within the Atomistic Line Graph Neural Network (ALIGNN) architecture. These features enable flexibility for systems with electrified interfaces, such as in electrocatalysts and achieves mean absolute error (MAE) of less than 0.1 eV. The machine learning model reported here achieves a strong generalization to left-out adsorbates (MAE = 0.074 eV) and leave-one-chemical-space-out (MAE = 0.08 eV) and completely left-out metals (MAE = 0.072 eV), confirming the robustness of the machine learning (ML) model in predicting ..epsilon..nkHSE.

36 MATERIALS SCIENCE↗

Machine Learning Correlation of Electron Micrographs and ToF-SIMS for the Analysis of Organic Biomarkers in Mudstone

The spatial distribution of organics in geological samples can be used to determine when and how these organics were incorporated into the host rock. Mass spectrometry (MS) imaging can rapidly collect a large amount of data, but ions produced are mixed without discrimination, resulting in complex mass spectra that can be difficult to interpret. Here, we apply unsupervised and supervised machine learning (ML) to help interpret spectra from time-of-flight-secondary ion mass spectrometry (ToF-SIMS) of an organic-carbon-rich mudstone of the Middle Jurassic of England (UK). It was previously shown that the presence of sterane molecular biomarkers in this sample can be detected via ToF-SIMS (Pasterski, M. J. et al., Astrobiology 2023, 23, 936). We use unsupervised ML on scanning electron microscopy–electron dispersive spectroscopy (SEM-EDS) measurements to define compositional categories based on differences in elemental abundances. We then test the ability of four ML algorithms─k-nearest neighbors (KNN), recursive partitioning and regressive trees (RPART), eXtreme gradient boost (XGBoost), and random forest (RF)─to classify the ToF-SIM spectra using (1) the categories assigned via SEM-EDS, (2) organic and inorganic labels assigned via SEM-EDS, and (3) the presence or absence of detectable steranes in ToF-SIMS spectra. In terms of predictive accuracy and balanced accuracy, KNN was the best performing model and RPART the worst. The feature importance, or the specific features of the ToF-SIM spectra used by the models to make classifications, cannot be determined for KNN, preventing posthoc model interpretation. Nevertheless, the feature importance extracted from the other models was useful for interpreting spectra. In conclusion, we determined that some of the organic ions used to classify biomarker containing spectra may be fragment ions derived from kerogen which is abundant in this mudstone sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Mass‐Conserving‐Perceptron for Machine‐Learning‐Based Modeling of Geoscientific Systems

Although decades of effort have been devoted to building Physical-Conceptual (PC) models for predicting the time-series evolution of geoscientific systems, recent work shows that Machine Learning (ML) based Gated Recurrent Neural Network technology can be used to develop models that are much more accurate. However, the difficulty of extracting physical understanding from ML-based models complicates their utility for enhancing scientific knowledge regarding system structure and function. Here, we propose a physically interpretable Mass-Conserving-Perceptron (MCP) as a way to bridge the gap between PC-based and ML-based modeling approaches. The MCP exploits the inherent isomorphism between the directed graph structures underlying both PC models and GRNNs to explicitly represent the mass-conserving nature of physical processes while enabling the functional nature of such processes to be directly learned (in an interpretable manner) from available data using off-the-shelf ML technology. As a proof of concept, we investigate the functional expressivity (capacity) of the MCP, explore its ability to parsimoniously represent the rainfall-runoff (RR) dynamics of the Leaf River Basin, and demonstrate its utility for scientific hypothesis testing. To conclude, we discuss extensions of the concept to enable ML-based physical-conceptual representation of the coupled nature of mass-energy-information flows through geoscientific systems.

58 GEOSCIENCES↗

Projecting Large Fires in the Western US With an Interpretable and Accurate Hybrid Machine Learning Method

More frequent and widespread large fires are occurring in the western United States (US), yet reliable methods for predicting these fires, particularly with extended lead times and a high spatial resolution, remain challenging. In this study, we proposed an interpretable and accurate hybrid machine learning (ML) model, that explicitly represented the controls of fuel flammability, fuel availability, and human suppression effects on fires. The model demonstrated notable accuracy with a F 1 -score of 0.846 ± 0.012, surpassing process-driven fire danger indices and four commonly used ML models by up to 40% and 9%, respectively. More importantly, the ML model showed remarkably higher interpretability relative to other ML models. Specifically, by demystifying the “black box” of each ML model using the explainable AI techniques, we identified substantial structural differences across ML fire models, even among those with similar accuracy. The relationships between fires and their drivers, identified by our model, were aligned closer with established fire physical principles. The ML structural discrepancy led to diverse fire predictions and our model predictions exhibited greater consistency with actual fire occurrence. With the highly interpretable and accurate model, we revealed the strong compound effects from multiple climate variables related to evaporative demand, energy release component, temperature, and wind speed, on the dynamics of large fires and megafires in the western US. Our findings highlight the importance of assessing the structural integrity of models in addition to their accuracy. They also underscore the critical need to address the rise in compound climate extremes linked to large wildfires.

54 ENVIRONMENTAL SCIENCES↗

Improving Low‐Cloud Fraction Prediction Through Machine Learning

Abstract In this study, we evaluated the performance of machine learning (ML) models (XGBoost) in predicting low‐cloud fraction (LCF), compared to two generations of the community atmospheric model (CAM5 and CAM6) and ERA5 reanalysis data, each having a different cloud scheme. ML models show a substantial enhancement in predicting LCF regarding root mean squared errors and correlation coefficients. The good performance is consistent across the full spectrums of atmospheric stability and large‐scale vertical velocity. Employing an explainable ML approach, we revealed the importance of including the amount of available moisture in ML models for representing spatiotemporal variations in LCF in the midlatitudes. Also, ML models demonstrated marked improvement in capturing the LCF variations during the stratocumulus‐to‐cumulus transition (SCT). This study suggests ML models' great potential to address the longstanding issues of “too few” low clouds and “too rapid” SCT in global climate models.

Geology↗

Empowering Machine Learning Forecasting of Labquake Using Event‐Based Features and Clustering Characteristics

Abstract Following recent advances of machine learning (ML), we present a novel approach to extract spatiotemporal seismo‐mechanical features from Acoustic Emission (AE) catalogs to empower ML‐based forecasting. The AE data were recorded during laboratory stick‐slip experiments on granite samples cut by rough faults. Based on the features computed for a past time window, a random forest (RF) classifier is used to forecast the occurrence of a large magnitude event ( M AE > 3.5) in the next time window. Event‐based features allow us to associate informative time‐space characteristics to each feature and nearest‐neighbor clustering analysis enables us to separate background and clustered seismicity and train individual models. The results show that the separation of AEs enhances the forecasting accuracy from 73.2% for the entire catalog up to 82.1% and 89.0% if background and clustered events are used separately. The presented new approach may be upscaled for applications to forecast tectonic earthquakes.

Karimpouli, Sadegh↗

Unsupervised Clustering of Microseismic Events and Focal Mechanism Analysis at the CO 2 Injection Site in Decatur, Illinois

Characterization of induced microseismicity at a carbon dioxide (CO 2 ) storage site is critical for preserving reservoir integrity and mitigating seismic hazards. We apply a multilevel machine learning (ML) approach that combines the nonnegative matrix factorization and hidden Markov model to extract spectral representations of microseismic events and cluster them to identify seismic patterns at the Illinois Basin-Decatur Project. Unlike traditional waveform correlation methods, this approach leverages spectral characteristics of first arrivals to improve event classification and detect previously undetected planes of weakness. By integrating ML-based clustering with focal mechanism analysis, we resolve small-scale fault structures that are below the detection limits of conventional seismic imaging. Our findings reveal temporal bursts of microseismicity associated with brittle failure, providing insights into the spatio-temporal evolution of fault reactivation during CO 2 injection. This approach enhances seismic monitoring capabilities at CO 2 injection sites by improving fault characterization beyond the resolution of standard geophysical surveys.

Willis, Rachel Marie [Sandia National Laboratories↗

Stable Simulation of the Community Atmosphere Model Using Machine‐Learning Physical Parameterization Trained With Experience Replay

In recent years, machine learning (ML) models have been used to improve physical parameterizations of general circulation models (GCMs). A significant challenge of integrating ML models into GCMs is the online instability when they are coupled for long‐term simulation. We present a new strategy that demonstrates robust online stability when the physical parameterization package of an atmospheric GCM is replaced by a deep ML model. The method uses experience replay with a multistep training scheme of the ML model in which the model's own output at the previous time step is used in the training. Predicted physics tendencies in the replay buffer with the most recent errors in the training iterations are reused, making the ML model learn from its own errors. The training method reduces the gap between the offline and online environments of the ML model. The method is used to train the ML model as the physical parameterization of the Community Atmosphere Model (CAM5) with training data from the Multi‐scale Modeling Framework high resolution simulations. Three 6‐year online simulations of the CAM5 are carried out by using the ML physics package. The simulated spatial distributions of precipitation, surface temperature and zonally averaged atmospheric fields demonstrate overall better accuracy than that of the standard CAM5 and benchmark model even without the use of additional physical constraints or tuning. This work is the first to demonstrate a solution to address the online instability problem in climate modeling with ML physics by using experience replay.

54 ENVIRONMENTAL SCIENCES↗

A Probabilistic Model for Global EMIC Wave Activity Using Van Allen Probes Observations

Electromagnetic ion cyclotron (EMIC) waves play a key role in radiation belt dynamics through resonant interactions. However, their low occurrence probability, high variability, and spatial intermittency pose challenges for accurate modeling. In this study, we present a machine learning (ML)-based global EMIC wave model built on the entire data set from the Van Allen Probes mission. To capture the distinct statistical characteristics of wave occurrence and amplitude, the model is separated into two modules: an occurrence model trained using ML techniques, and a wave amplitude model sampled from observed probability distributions. The input parameters are limited to real-time or predictable variables to ensure practical applicability. Our model shows strong performance across the entire test set and demonstrates improved predictive capability over a baseline random occurrence model, particularly during quiet geomagnetic conditions. Evaluation during both quiet and active periods confirms the model's ability to represent the clustered and intermittent nature of EMIC wave activity. Furthermore, the model provides global estimates of wave power, enabling integration with radiation belt electron data and showing signatures consistent with wave-induced scattering. We found a good correlation between the global wave activity from the model and relativistic electron observation by Van Allen Probes, regardless of the availability of in situ wave observations. The modular structure of the model also allows for straightforward expansion for additional wave properties, such as wave frequency, which can be modeled independently. This flexible, event-sensitive approach offers a promising framework for data-driven radiation belt simulations and space weather applications.

79 ASTRONOMY AND ASTROPHYSICS↗

The Value of Forecasters‐in‐the‐Loop in Real‐Time Flood Forecasting in the Age of Machine Learning

Machine learning (ML) applications in hydrological forecasting are increasingly prevalent and show great potential. However, many previous studies have only evaluated performance through reanalysis or retrospective simulations compared to simplified baselines. This study provides the first assessment of ML performance against actual operational forecasting systems operated by the California Nevada River Forecast Center (CNRFC), which combines the Community Hydrologic Prediction System (CHPS) with forecasters-in-the-loop. Results demonstrate that forecasters-in-the-loop systems consistently outperform ML models in both general forecasts and flood alerting across lead times up to 96 hr, even when ML models use observed forcings, while CNRFC operational process relies on biased weather forecasts. Our analysis reveals that forecaster expertise maintains forecast reliability despite inaccurate precipitation inputs, with human-guided systems showing superior performance degradation characteristics at extended lead times. These findings highlight the irreplaceable value of human expertise in operational forecasting and caution against overstating current ML capabilities in real-world applications.

Tran, Vinh Ngoc [Univ. of Michigan, Ann Arbor, MI ↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

A Machine Learning Framework for Predicting Microphysical Properties of Ice Crystals From Cloud Particle Imagery

The microphysical properties of ice crystals are important because they significantly alter the radiative properties and spatiotemporal distributions of clouds, which in turn strongly affect Earth's climate. However, it is challenging to measure key properties of ice crystals, such as mass or morphological features. Here, we present a proof-of-concept framework for predicting three-dimensional (3D) microphysical properties of ice crystals from in situ two-dimensional (2D) imagery. First, we computationally generated synthetic ice crystals using 3D modeling software along with geometric parameters estimated from the 2021 Ice Cryo-Encapsulation Balloon (ICEBall) field campaign. Then, we used synthetic crystals to train machine learning (ML) models to predict effective density ($ρ_e$), effective surface area ($A_e$), and number of bullets ($N_b$) from synthetic rosette imagery. On unseen synthetic images, our ML models accurately predicted ice crystal properties. ResNet-18 performed best, achieving $R^2$ values of 0.99 and 0.98 for $ρ_e$ and $A_e$, respectively, and MAE of 0.10 for mathematical equation in single view tasks. Stereo view ResNet-18 further reduced RMSE by 40% for $ρ_e$ and $A_e$ and reduced MAE by 0.08 for $N_b$. This work provides a novel ML-driven framework for estimating ice microphysical properties from in situ imagery, which will allow for downstream constraints on microphysical parameterizations, such as the mass-size relationship.

Ko, J. [Columbia Univ., New York, NY (United State↗

A Decadal Hybrid GCM Simulation Using Deep‐Learning‐Based Cloud and Convection Parameterization Generalized to a Warm Climate

A critical challenge for machine‐learning (ML) parameterization in global climate models (GCMs) is to achieve stable, accurate simulations under climates not seen during training. Previous studies have demonstrated promising offline performance and year‐long online stability in aquaplanet simulations but have encountered difficulties in real geography and under climate warming. Here we report that a GCM with real geography configuration using neural‐network‐based cloud and convection parameterization, trained exclusively with present‐day climate data, successfully performs a stable, decade‐long simulation of a warm climate with +4 K sea surface temperature (SST). The neural network (NN) is based on Han et al. (2023, https://doi.org/10.1029/2022ms003508 ) with additional inputs. The simulation captures the global precipitation distribution, surface temperatures, vertical atmospheric structures, and extreme precipitation very well, closely matching simulations from both the superparameterized CAM (SPCAM) and the conventional CAM5 in the warm climate without accuracy degradation compared to those in the baseline climate. Moreover, it produces a climate response to +4 K SST in atmospheric thermodynamic states and circulations similar to those from SPCAM and CAM5. Prognostic ablation tests on NN input variables show that the NN without convective memory as input suffers from numerical instability, and the NN without considering radiative variables and land fraction as input, or with reduced training samples produce less accurate results. To our knowledge, this is the first time an ML parameterization successfully achieves online extrapolation to a warm climate without using additional warm‐climate data for training. It demonstrates the potential of ML‐driven parameterizations for credible long‐term climate projections.

Atmosphere model↗

Perspectives on Systematic Cloud Microphysics Scheme Development With Machine Learning

Cloud microphysics—the collection of processes that govern the small‐scale formation, evolution, and interactions of liquid droplets and ice crystals in clouds and precipitation—remains a major source of uncertainty in weather and climate models. Although too small in scale to be explicitly resolved in any large‐eddy simulation, weather, or climate model, the representation of cloud microphysical processes has significant impact at the climate scale. Current microphysical schemes are limited by both parametric uncertainty, linked to uncertainty in physical parameter values, and structural uncertainty, arising from incomplete physical understanding of the processes at play or approximations made for computational efficiency. Recent advances in the application of machine learning (ML) to the physical sciences show significant potential for minimizing these limitations by leveraging high‐fidelity simulations and observations. Here we outline the challenges that must be addressed to apply ML toward cloud microphysics scheme development. This perspectives paper synthesizes recent progress in using data‐driven methods, including ML, to improve cloud microphysics parameterizations and highlights opportunities to address key uncertainties. We discuss the roles of aleatoric (irreducible, or statistical) and epistemic (reducible, or systematic) errors in contributing to microphysics parameterization uncertainty. ML can leverage observations to improve microphysical schemes via bottom‐up and top‐down constraints. Methods such as differentiable programming and ML‐enhanced sampling strategies and the creation of large scale benchmark data sets promise to bridge the gap between observations and models and to improve the consistency of cloud microphysical representation across temporal and spatial scales.

Lamb, Kara D. [Columbia Univ., New York, NY (Unite↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

Hydrology in the Age of Artificial Intelligence: From Fragmentation to Coherent Terrestrial Hydrosphere Science

The rapid rise of machine learning (ML) in hydrology has prompted debate about the discipline's scientific relevance. While ML often outperforms traditional models in streamflow prediction, we argue that this reflects a deeper limitation: persistent fragmentation of hydrological science itself. Narrow focus on isolated components has hindered the development of coherent, scale‐relevant understanding of the integrated terrestrial hydrosphere. This is illustrated, for example, by widely divergent estimates of groundwater–streamflow interactions and of water balance‐implied ongoing storage changes. We argue that hydrology's future lies not in choosing between ML and physics, but in integrating data‐driven and process‐based approaches to advance consistent, realistic, and societally relevant understanding of the terrestrial hydrosphere and its multifaceted roles in the Earth System.

Painter, Scott L. [Oak Ridge National Laboratory (↗