Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning (ML)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

A Probabilistic Model for Global EMIC Wave Activity Using Van Allen Probes Observations

Electromagnetic ion cyclotron (EMIC) waves play a key role in radiation belt dynamics through resonant interactions. However, their low occurrence probability, high variability, and spatial intermittency pose challenges for accurate modeling. In this study, we present a machine learning (ML)-based global EMIC wave model built on the entire data set from the Van Allen Probes mission. To capture the distinct statistical characteristics of wave occurrence and amplitude, the model is separated into two modules: an occurrence model trained using ML techniques, and a wave amplitude model sampled from observed probability distributions. The input parameters are limited to real-time or predictable variables to ensure practical applicability. Our model shows strong performance across the entire test set and demonstrates improved predictive capability over a baseline random occurrence model, particularly during quiet geomagnetic conditions. Evaluation during both quiet and active periods confirms the model's ability to represent the clustered and intermittent nature of EMIC wave activity. Furthermore, the model provides global estimates of wave power, enabling integration with radiation belt electron data and showing signatures consistent with wave-induced scattering. We found a good correlation between the global wave activity from the model and relativistic electron observation by Van Allen Probes, regardless of the availability of in situ wave observations. The modular structure of the model also allows for straightforward expansion for additional wave properties, such as wave frequency, which can be modeled independently. This flexible, event-sensitive approach offers a promising framework for data-driven radiation belt simulations and space weather applications.

79 ASTRONOMY AND ASTROPHYSICS↗

The Value of Forecasters‐in‐the‐Loop in Real‐Time Flood Forecasting in the Age of Machine Learning

Machine learning (ML) applications in hydrological forecasting are increasingly prevalent and show great potential. However, many previous studies have only evaluated performance through reanalysis or retrospective simulations compared to simplified baselines. This study provides the first assessment of ML performance against actual operational forecasting systems operated by the California Nevada River Forecast Center (CNRFC), which combines the Community Hydrologic Prediction System (CHPS) with forecasters-in-the-loop. Results demonstrate that forecasters-in-the-loop systems consistently outperform ML models in both general forecasts and flood alerting across lead times up to 96 hr, even when ML models use observed forcings, while CNRFC operational process relies on biased weather forecasts. Our analysis reveals that forecaster expertise maintains forecast reliability despite inaccurate precipitation inputs, with human-guided systems showing superior performance degradation characteristics at extended lead times. These findings highlight the irreplaceable value of human expertise in operational forecasting and caution against overstating current ML capabilities in real-world applications.

Tran, Vinh Ngoc [Univ. of Michigan, Ann Arbor, MI ↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

A Machine Learning Framework for Predicting Microphysical Properties of Ice Crystals From Cloud Particle Imagery

The microphysical properties of ice crystals are important because they significantly alter the radiative properties and spatiotemporal distributions of clouds, which in turn strongly affect Earth's climate. However, it is challenging to measure key properties of ice crystals, such as mass or morphological features. Here, we present a proof-of-concept framework for predicting three-dimensional (3D) microphysical properties of ice crystals from in situ two-dimensional (2D) imagery. First, we computationally generated synthetic ice crystals using 3D modeling software along with geometric parameters estimated from the 2021 Ice Cryo-Encapsulation Balloon (ICEBall) field campaign. Then, we used synthetic crystals to train machine learning (ML) models to predict effective density ($ρ_e$), effective surface area ($A_e$), and number of bullets ($N_b$) from synthetic rosette imagery. On unseen synthetic images, our ML models accurately predicted ice crystal properties. ResNet-18 performed best, achieving $R^2$ values of 0.99 and 0.98 for $ρ_e$ and $A_e$, respectively, and MAE of 0.10 for mathematical equation in single view tasks. Stereo view ResNet-18 further reduced RMSE by 40% for $ρ_e$ and $A_e$ and reduced MAE by 0.08 for $N_b$. This work provides a novel ML-driven framework for estimating ice microphysical properties from in situ imagery, which will allow for downstream constraints on microphysical parameterizations, such as the mass-size relationship.

Ko, J. [Columbia Univ., New York, NY (United State↗

A Decadal Hybrid GCM Simulation Using Deep‐Learning‐Based Cloud and Convection Parameterization Generalized to a Warm Climate

A critical challenge for machine‐learning (ML) parameterization in global climate models (GCMs) is to achieve stable, accurate simulations under climates not seen during training. Previous studies have demonstrated promising offline performance and year‐long online stability in aquaplanet simulations but have encountered difficulties in real geography and under climate warming. Here we report that a GCM with real geography configuration using neural‐network‐based cloud and convection parameterization, trained exclusively with present‐day climate data, successfully performs a stable, decade‐long simulation of a warm climate with +4 K sea surface temperature (SST). The neural network (NN) is based on Han et al. (2023, https://doi.org/10.1029/2022ms003508 ) with additional inputs. The simulation captures the global precipitation distribution, surface temperatures, vertical atmospheric structures, and extreme precipitation very well, closely matching simulations from both the superparameterized CAM (SPCAM) and the conventional CAM5 in the warm climate without accuracy degradation compared to those in the baseline climate. Moreover, it produces a climate response to +4 K SST in atmospheric thermodynamic states and circulations similar to those from SPCAM and CAM5. Prognostic ablation tests on NN input variables show that the NN without convective memory as input suffers from numerical instability, and the NN without considering radiative variables and land fraction as input, or with reduced training samples produce less accurate results. To our knowledge, this is the first time an ML parameterization successfully achieves online extrapolation to a warm climate without using additional warm‐climate data for training. It demonstrates the potential of ML‐driven parameterizations for credible long‐term climate projections.

Atmosphere model↗

Perspectives on Systematic Cloud Microphysics Scheme Development With Machine Learning

Cloud microphysics—the collection of processes that govern the small‐scale formation, evolution, and interactions of liquid droplets and ice crystals in clouds and precipitation—remains a major source of uncertainty in weather and climate models. Although too small in scale to be explicitly resolved in any large‐eddy simulation, weather, or climate model, the representation of cloud microphysical processes has significant impact at the climate scale. Current microphysical schemes are limited by both parametric uncertainty, linked to uncertainty in physical parameter values, and structural uncertainty, arising from incomplete physical understanding of the processes at play or approximations made for computational efficiency. Recent advances in the application of machine learning (ML) to the physical sciences show significant potential for minimizing these limitations by leveraging high‐fidelity simulations and observations. Here we outline the challenges that must be addressed to apply ML toward cloud microphysics scheme development. This perspectives paper synthesizes recent progress in using data‐driven methods, including ML, to improve cloud microphysics parameterizations and highlights opportunities to address key uncertainties. We discuss the roles of aleatoric (irreducible, or statistical) and epistemic (reducible, or systematic) errors in contributing to microphysics parameterization uncertainty. ML can leverage observations to improve microphysical schemes via bottom‐up and top‐down constraints. Methods such as differentiable programming and ML‐enhanced sampling strategies and the creation of large scale benchmark data sets promise to bridge the gap between observations and models and to improve the consistency of cloud microphysical representation across temporal and spatial scales.

Lamb, Kara D. [Columbia Univ., New York, NY (Unite↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

Hydrology in the Age of Artificial Intelligence: From Fragmentation to Coherent Terrestrial Hydrosphere Science

The rapid rise of machine learning (ML) in hydrology has prompted debate about the discipline's scientific relevance. While ML often outperforms traditional models in streamflow prediction, we argue that this reflects a deeper limitation: persistent fragmentation of hydrological science itself. Narrow focus on isolated components has hindered the development of coherent, scale‐relevant understanding of the integrated terrestrial hydrosphere. This is illustrated, for example, by widely divergent estimates of groundwater–streamflow interactions and of water balance‐implied ongoing storage changes. We argue that hydrology's future lies not in choosing between ML and physics, but in integrating data‐driven and process‐based approaches to advance consistent, realistic, and societally relevant understanding of the terrestrial hydrosphere and its multifaceted roles in the Earth System.

Painter, Scott L. [Oak Ridge National Laboratory (↗

Analytical ab initio hessian from a deep learning potential for transition state optimization

Identifying transition states—saddle points on the potential energy surface connecting reactant and product minima—is central to predicting kinetic barriers and understanding chemical reaction mechanisms. In this work, we train a fully differentiable equivariant neural network potential, NewtonNet, on thousands of organic reactions and derive the analytical Hessians. By reducing the computational cost by several orders of magnitude relative to the density functional theory (DFT) ab initio source, we can afford to use the learned Hessians at every step for the saddle point optimizations. We show that the full machine learned (ML) Hessian robustly finds the transition states of 240 unseen organic reactions, even when the quality of the initial guess structures are degraded, while reducing the number of optimization steps to convergence by 2–3× compared to the quasi-Newton DFT and ML methods. All data generation, NewtonNet model, and ML transition state finding methods are available in an automated workflow.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Model-free estimation of completeness, uncertainties, and outliers in atomistic machine learning using information theory

Abstract An accurate description of information is relevant for a range of problems in atomistic machine learning (ML), such as crafting training sets, performing uncertainty quantification (UQ), or extracting physical insights from large datasets. However, atomistic ML often relies on unsupervised learning or model predictions to analyze information contents from simulation or training data. Here, we introduce a theoretical framework that provides a rigorous, model-free tool to quantify information contents in atomistic simulations. We demonstrate that the information entropy of a distribution of atom-centered environments explains known heuristics in ML potential developments, from training set sizes to dataset optimality. Using this tool, we propose a model-free UQ method that reliably predicts epistemic uncertainty and detects out-of-distribution samples, including rare events in systems such as nucleation. This method provides a general tool for data-driven atomistic modeling and combines efforts in ML, simulations, and physical explainability.

36 MATERIALS SCIENCE↗

Optical neural engine for solving scientific partial differential equations

Abstract Solving partial differential equations (PDEs) is the cornerstone of scientific research and development. Data-driven machine learning (ML) approaches are emerging to accelerate time-consuming and computation-intensive numerical simulations of PDEs. Although optical systems offer high-throughput and energy-efficient ML hardware, their demonstration for solving PDEs is limited. Here, we present an optical neural engine (ONE) architecture combining diffractive optical neural networks for Fourier space processing and optical crossbar structures for real space processing to solve time-dependent and time-independent PDEs in diverse disciplines, including Darcy flow equation, the magnetostatic Poisson’s equation in demagnetization, the Navier-Stokes equation in incompressible fluid, Maxwell’s equations in nanophotonic metasurfaces, and coupled PDEs in a multiphysics system. We numerically and experimentally demonstrate the capability of the ONE architecture, which not only leverages the advantages of high-performance dual-space processing for outperforming traditional PDE solvers and being comparable with state-of-the-art ML models but also can be implemented using optical computing hardware with unique features of low-energy and highly parallel constant-time processing irrespective of model scales and real-time reconfigurability for tackling multiple tasks with the same architecture. The demonstrated architecture offers a versatile and powerful platform for large-scale scientific and engineering computations.

Tang, Yingheng (ORCID:0009000153622546)↗

Machine learning reveals genes impacting oxidative stress resistance across yeasts

Reactive oxygen species (ROS) are highly reactive molecules encountered by yeasts during routine metabolism and during interactions with other organisms, including host infection. Here, we characterize the variation in resistance to the ROS-inducing compound tert -butyl hydroperoxide across the ancient yeast subphylum Saccharomycotina and use machine learning (ML) to identify gene families whose sizes are predictive of ROS resistance. The most predictive features are enriched in gene families related to cell wall organization and include two reductase gene families. We estimate the quantitative contributions of features to each species’ classification to guide experimental validation and show that overexpression of the old yellow enzyme (OYE) reductase increases ROS resistance in Kluyveromyces lactis , while Saccharomyces cerevisiae mutants lacking multiple mannosyltransferase-encoding genes are hypersensitive to ROS. Altogether, this work provides a framework for how ML can uncover genetic mechanisms underlying trait variation across diverse species and inform trait manipulation for clinical and biotechnological applications.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning with plasma plume image sequences for anomaly detection and prediction of growth kinetics during pulsed laser deposition

Abstract Materials synthesis platforms that are designed for autonomous experimentation are capable of collecting multimodal diagnostic data that can be utilized for feedback to optimize material properties. Pulsed laser deposition (PLD) is emerging as a viable autonomous synthesis tool, and so the need arises to develop machine learning (ML) techniques that are capable of extracting information from in situ diagnostics. Here, we demonstrate that intensified-CCD image sequences of the plasma plume generated during PLD can be used for anomaly detection and the prediction of thin film growth kinetics. We develop multi-output (2 + 1)D convolutional neural network regression models that extract deep features from plume dynamics that not only correlate with the measured chamber pressure and incident laser energy, but more importantly, predict parameters of an auto-catalytic film growth model derived from in situ laser reflectivity experiments. Our results demonstrate how ML with in situ plume diagnostics data in PLD can be utilized to maintain deposition conditions in an optimal regime. Further, the predictive capabilities of plume dynamics on the kinetics of film growth or other film properties prior to deposition provides a means for rapid pre-screening of growth conditions for the non-expert, which promises to accelerate materials optimization with PLD.

36 MATERIALS SCIENCE↗

Electronic structure prediction of multi-million atom systems through uncertainty quantification enabled transfer learning

The ground state electron density — obtainable using Kohn-Sham Density Functional Theory (KS-DFT) simulations — contains a wealth of material information, making its prediction via machine learning (ML) models attractive. However, the computational expense of KS-DFT scales cubically with system size which tends to stymie training data generation, making it difficult to develop quantifiably accurate ML models that are applicable across many scales and system configurations. Here, we address this fundamental challenge by employing transfer learning to leverage the multi-scale nature of the training data, while comprehensively sampling system configurations using thermalization. Our ML models are less reliant on heuristics, and being based on Bayesian neural networks, enable uncertainty quantification. We show that our models incur significantly lower data generation costs while allowing confident — and when verifiable, accurate — predictions for a wide variety of bulk systems well beyond training, including systems with defects, different alloy compositions, and at multi-million-atom scales. Moreover, such predictions can be carried out using only modest computational resources.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerating phase field simulations through a hybrid adaptive Fourier neural operator with U-net backbone

Prolonged contact between a corrosive liquid and metal alloys can cause progressive dealloying. For one such process as liquid-metal dealloying (LMD), phase field models have been developed to understand the mechanisms leading to complex morphologies. However, the LMD governing equations in these models often involve coupled non-linear partial differential equations (PDE), which are challenging to solve numerically. In particular, numerical stiffness in the PDEs requires an extremely refined time step size (on the order of 10 -12 s or smaller). This computational bottleneck is especially problematic when running LMD simulation until a late time horizon is required. This motivates the development of surrogate models capable of leaping forward in time, by skipping several consecutive time steps at-once. In this paper, we propose a U-shaped adaptive Fourier neural operator (U-AFNO), a machine learning (ML) based model inspired by recent advances in neural operator learning. U-AFNO employs U-Nets for extracting and reconstructing local features within the physical fields, and passes the latent space through a vision transformer (ViT) implemented in the Fourier space (AFNO). We use U-AFNOs to learn the dynamics of mapping the field at a current time step into a later time step. We also identify global quantities of interest (QoI) describing the corrosion process (e.g., the deformation of the liquid-metal interface, lost metal, etc.) and show that our proposed U-AFNO model is able to accurately predict the field dynamics, in spite of the chaotic nature of LMD. Most notably, our model reproduces the key microstructure statistics and QoIs with a level of accuracy on par with the high-fidelity numerical solver, while achieving a significant 11, 200 × speed-up on a high-resolution grid when comparing the computational expense per time step. Finally, we also investigate the opportunity of using hybrid simulations, in which we alternate forward leaps in time using the U-AFNO with high-fidelity time stepping. We demonstrate that while advantageous for some surrogate model design choices, our proposed U-AFNO model in fully auto-regressive settings consistently outperforms hybrid schemes.

36 MATERIALS SCIENCE↗

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca↗

Attention-based functional-group coarse-graining: a deep learning framework for molecular prediction and design

Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.

Han, Ming [Univ. of Chicago, IL (United States)]↗

Electronic structure prediction of medium and high entropy alloys across composition space

We propose machine learning (ML) models to predict the electron density — the fundamental unknown of a material’s ground state — across the composition space of concentrated alloys. From this, other physical properties can be inferred, enabling accelerated exploration. A significant challenge is that the number of descriptors and sampled compositions required for accurate prediction grows rapidly with species. To address this, we employ Bayesian Active Learning (AL), which minimizes training data requirements by leveraging uncertainty quantification capabilities of Bayesian Neural Networks. Compared to the strategic tessellation of the composition space, Bayesian-AL reduces the number of training data points by a factor of 2.5 for ternary (SiGeSn) and 1.7 for quaternary (CrFeCoNi) systems. We also introduce easy-to-optimize, body-attached-frame descriptors, which respect physical symmetries while keeping descriptor-vector size nearly constant as alloy complexity increases. Our ML models demonstrate high accuracy and generalizability in predicting both electron density and energy across composition space.

materials science↗