Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning (ML)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Hyper‐Local Temperature Prediction Using Detailed Urban Climate Informatics

The accurate modeling of urban microclimate is a challenging task given the high surface heterogeneity of urban land cover and the vertical structure of street morphology. Recent years have witnessed significant efforts in numerical modeling and data collection of the urban environment. Nonetheless, it is difficult for the physical‐based models to fully utilize the high‐resolution data under the constraints of computing resources. The advancement in machine learning (ML) techniques offers the computational strength to handle the massive volume of data. In this study, we proposed a modeling framework that uses ML approach to estimate point‐scale street‐level air temperature from the urban‐resolving meso‐scale climate model and a suite of hyper‐resolution urban geospatial data sets, including three‐dimensional urban morphology, parcel‐level land use inventory, and weather observations from a sensor network. We implemented this approach in the City of Chicago as a case study to demonstrate the capability of the framework. The proposed approach vastly improves the resolution of temperature predictions in cities, which will help the city with walkability, drivability, and heat‐related behavioral studies. Moreover, we tested the model's reliability on out‐of‐sample locations to investigate the modeling uncertainties and the application potentials to the other areas. This study aims to gain insights into next‐gen urban climate modeling and guide the observation efforts in cities to build the strength for the holistic understanding of urban microclimate dynamics.

54 ENVIRONMENTAL SCIENCES↗

A Mass‐Conserving‐Perceptron for Machine‐Learning‐Based Modeling of Geoscientific Systems

Although decades of effort have been devoted to building Physical-Conceptual (PC) models for predicting the time-series evolution of geoscientific systems, recent work shows that Machine Learning (ML) based Gated Recurrent Neural Network technology can be used to develop models that are much more accurate. However, the difficulty of extracting physical understanding from ML-based models complicates their utility for enhancing scientific knowledge regarding system structure and function. Here, we propose a physically interpretable Mass-Conserving-Perceptron (MCP) as a way to bridge the gap between PC-based and ML-based modeling approaches. The MCP exploits the inherent isomorphism between the directed graph structures underlying both PC models and GRNNs to explicitly represent the mass-conserving nature of physical processes while enabling the functional nature of such processes to be directly learned (in an interpretable manner) from available data using off-the-shelf ML technology. As a proof of concept, we investigate the functional expressivity (capacity) of the MCP, explore its ability to parsimoniously represent the rainfall-runoff (RR) dynamics of the Leaf River Basin, and demonstrate its utility for scientific hypothesis testing. To conclude, we discuss extensions of the concept to enable ML-based physical-conceptual representation of the coupled nature of mass-energy-information flows through geoscientific systems.

58 GEOSCIENCES↗

Projecting Large Fires in the Western US With an Interpretable and Accurate Hybrid Machine Learning Method

More frequent and widespread large fires are occurring in the western United States (US), yet reliable methods for predicting these fires, particularly with extended lead times and a high spatial resolution, remain challenging. In this study, we proposed an interpretable and accurate hybrid machine learning (ML) model, that explicitly represented the controls of fuel flammability, fuel availability, and human suppression effects on fires. The model demonstrated notable accuracy with a F 1 -score of 0.846 ± 0.012, surpassing process-driven fire danger indices and four commonly used ML models by up to 40% and 9%, respectively. More importantly, the ML model showed remarkably higher interpretability relative to other ML models. Specifically, by demystifying the “black box” of each ML model using the explainable AI techniques, we identified substantial structural differences across ML fire models, even among those with similar accuracy. The relationships between fires and their drivers, identified by our model, were aligned closer with established fire physical principles. The ML structural discrepancy led to diverse fire predictions and our model predictions exhibited greater consistency with actual fire occurrence. With the highly interpretable and accurate model, we revealed the strong compound effects from multiple climate variables related to evaporative demand, energy release component, temperature, and wind speed, on the dynamics of large fires and megafires in the western US. Our findings highlight the importance of assessing the structural integrity of models in addition to their accuracy. They also underscore the critical need to address the rise in compound climate extremes linked to large wildfires.

54 ENVIRONMENTAL SCIENCES↗

Improving Low‐Cloud Fraction Prediction Through Machine Learning

Abstract In this study, we evaluated the performance of machine learning (ML) models (XGBoost) in predicting low‐cloud fraction (LCF), compared to two generations of the community atmospheric model (CAM5 and CAM6) and ERA5 reanalysis data, each having a different cloud scheme. ML models show a substantial enhancement in predicting LCF regarding root mean squared errors and correlation coefficients. The good performance is consistent across the full spectrums of atmospheric stability and large‐scale vertical velocity. Employing an explainable ML approach, we revealed the importance of including the amount of available moisture in ML models for representing spatiotemporal variations in LCF in the midlatitudes. Also, ML models demonstrated marked improvement in capturing the LCF variations during the stratocumulus‐to‐cumulus transition (SCT). This study suggests ML models' great potential to address the longstanding issues of “too few” low clouds and “too rapid” SCT in global climate models.

Geology↗

Empowering Machine Learning Forecasting of Labquake Using Event‐Based Features and Clustering Characteristics

Abstract Following recent advances of machine learning (ML), we present a novel approach to extract spatiotemporal seismo‐mechanical features from Acoustic Emission (AE) catalogs to empower ML‐based forecasting. The AE data were recorded during laboratory stick‐slip experiments on granite samples cut by rough faults. Based on the features computed for a past time window, a random forest (RF) classifier is used to forecast the occurrence of a large magnitude event ( M AE > 3.5) in the next time window. Event‐based features allow us to associate informative time‐space characteristics to each feature and nearest‐neighbor clustering analysis enables us to separate background and clustered seismicity and train individual models. The results show that the separation of AEs enhances the forecasting accuracy from 73.2% for the entire catalog up to 82.1% and 89.0% if background and clustered events are used separately. The presented new approach may be upscaled for applications to forecast tectonic earthquakes.

Karimpouli, Sadegh↗

Unsupervised Clustering of Microseismic Events and Focal Mechanism Analysis at the CO 2 Injection Site in Decatur, Illinois

Characterization of induced microseismicity at a carbon dioxide (CO 2 ) storage site is critical for preserving reservoir integrity and mitigating seismic hazards. We apply a multilevel machine learning (ML) approach that combines the nonnegative matrix factorization and hidden Markov model to extract spectral representations of microseismic events and cluster them to identify seismic patterns at the Illinois Basin-Decatur Project. Unlike traditional waveform correlation methods, this approach leverages spectral characteristics of first arrivals to improve event classification and detect previously undetected planes of weakness. By integrating ML-based clustering with focal mechanism analysis, we resolve small-scale fault structures that are below the detection limits of conventional seismic imaging. Our findings reveal temporal bursts of microseismicity associated with brittle failure, providing insights into the spatio-temporal evolution of fault reactivation during CO 2 injection. This approach enhances seismic monitoring capabilities at CO 2 injection sites by improving fault characterization beyond the resolution of standard geophysical surveys.

Willis, Rachel Marie [Sandia National Laboratories↗

Stable Simulation of the Community Atmosphere Model Using Machine‐Learning Physical Parameterization Trained With Experience Replay

In recent years, machine learning (ML) models have been used to improve physical parameterizations of general circulation models (GCMs). A significant challenge of integrating ML models into GCMs is the online instability when they are coupled for long‐term simulation. We present a new strategy that demonstrates robust online stability when the physical parameterization package of an atmospheric GCM is replaced by a deep ML model. The method uses experience replay with a multistep training scheme of the ML model in which the model's own output at the previous time step is used in the training. Predicted physics tendencies in the replay buffer with the most recent errors in the training iterations are reused, making the ML model learn from its own errors. The training method reduces the gap between the offline and online environments of the ML model. The method is used to train the ML model as the physical parameterization of the Community Atmosphere Model (CAM5) with training data from the Multi‐scale Modeling Framework high resolution simulations. Three 6‐year online simulations of the CAM5 are carried out by using the ML physics package. The simulated spatial distributions of precipitation, surface temperature and zonally averaged atmospheric fields demonstrate overall better accuracy than that of the standard CAM5 and benchmark model even without the use of additional physical constraints or tuning. This work is the first to demonstrate a solution to address the online instability problem in climate modeling with ML physics by using experience replay.

54 ENVIRONMENTAL SCIENCES↗

A Probabilistic Model for Global EMIC Wave Activity Using Van Allen Probes Observations

Electromagnetic ion cyclotron (EMIC) waves play a key role in radiation belt dynamics through resonant interactions. However, their low occurrence probability, high variability, and spatial intermittency pose challenges for accurate modeling. In this study, we present a machine learning (ML)-based global EMIC wave model built on the entire data set from the Van Allen Probes mission. To capture the distinct statistical characteristics of wave occurrence and amplitude, the model is separated into two modules: an occurrence model trained using ML techniques, and a wave amplitude model sampled from observed probability distributions. The input parameters are limited to real-time or predictable variables to ensure practical applicability. Our model shows strong performance across the entire test set and demonstrates improved predictive capability over a baseline random occurrence model, particularly during quiet geomagnetic conditions. Evaluation during both quiet and active periods confirms the model's ability to represent the clustered and intermittent nature of EMIC wave activity. Furthermore, the model provides global estimates of wave power, enabling integration with radiation belt electron data and showing signatures consistent with wave-induced scattering. We found a good correlation between the global wave activity from the model and relativistic electron observation by Van Allen Probes, regardless of the availability of in situ wave observations. The modular structure of the model also allows for straightforward expansion for additional wave properties, such as wave frequency, which can be modeled independently. This flexible, event-sensitive approach offers a promising framework for data-driven radiation belt simulations and space weather applications.

79 ASTRONOMY AND ASTROPHYSICS↗

The Value of Forecasters‐in‐the‐Loop in Real‐Time Flood Forecasting in the Age of Machine Learning

Machine learning (ML) applications in hydrological forecasting are increasingly prevalent and show great potential. However, many previous studies have only evaluated performance through reanalysis or retrospective simulations compared to simplified baselines. This study provides the first assessment of ML performance against actual operational forecasting systems operated by the California Nevada River Forecast Center (CNRFC), which combines the Community Hydrologic Prediction System (CHPS) with forecasters-in-the-loop. Results demonstrate that forecasters-in-the-loop systems consistently outperform ML models in both general forecasts and flood alerting across lead times up to 96 hr, even when ML models use observed forcings, while CNRFC operational process relies on biased weather forecasts. Our analysis reveals that forecaster expertise maintains forecast reliability despite inaccurate precipitation inputs, with human-guided systems showing superior performance degradation characteristics at extended lead times. These findings highlight the irreplaceable value of human expertise in operational forecasting and caution against overstating current ML capabilities in real-world applications.

Tran, Vinh Ngoc [Univ. of Michigan, Ann Arbor, MI ↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

A Machine Learning Framework for Predicting Microphysical Properties of Ice Crystals From Cloud Particle Imagery

The microphysical properties of ice crystals are important because they significantly alter the radiative properties and spatiotemporal distributions of clouds, which in turn strongly affect Earth's climate. However, it is challenging to measure key properties of ice crystals, such as mass or morphological features. Here, we present a proof-of-concept framework for predicting three-dimensional (3D) microphysical properties of ice crystals from in situ two-dimensional (2D) imagery. First, we computationally generated synthetic ice crystals using 3D modeling software along with geometric parameters estimated from the 2021 Ice Cryo-Encapsulation Balloon (ICEBall) field campaign. Then, we used synthetic crystals to train machine learning (ML) models to predict effective density ($ρ_e$), effective surface area ($A_e$), and number of bullets ($N_b$) from synthetic rosette imagery. On unseen synthetic images, our ML models accurately predicted ice crystal properties. ResNet-18 performed best, achieving $R^2$ values of 0.99 and 0.98 for $ρ_e$ and $A_e$, respectively, and MAE of 0.10 for mathematical equation in single view tasks. Stereo view ResNet-18 further reduced RMSE by 40% for $ρ_e$ and $A_e$ and reduced MAE by 0.08 for $N_b$. This work provides a novel ML-driven framework for estimating ice microphysical properties from in situ imagery, which will allow for downstream constraints on microphysical parameterizations, such as the mass-size relationship.

Ko, J. [Columbia Univ., New York, NY (United State↗

A Decadal Hybrid GCM Simulation Using Deep‐Learning‐Based Cloud and Convection Parameterization Generalized to a Warm Climate

A critical challenge for machine‐learning (ML) parameterization in global climate models (GCMs) is to achieve stable, accurate simulations under climates not seen during training. Previous studies have demonstrated promising offline performance and year‐long online stability in aquaplanet simulations but have encountered difficulties in real geography and under climate warming. Here we report that a GCM with real geography configuration using neural‐network‐based cloud and convection parameterization, trained exclusively with present‐day climate data, successfully performs a stable, decade‐long simulation of a warm climate with +4 K sea surface temperature (SST). The neural network (NN) is based on Han et al. (2023, https://doi.org/10.1029/2022ms003508 ) with additional inputs. The simulation captures the global precipitation distribution, surface temperatures, vertical atmospheric structures, and extreme precipitation very well, closely matching simulations from both the superparameterized CAM (SPCAM) and the conventional CAM5 in the warm climate without accuracy degradation compared to those in the baseline climate. Moreover, it produces a climate response to +4 K SST in atmospheric thermodynamic states and circulations similar to those from SPCAM and CAM5. Prognostic ablation tests on NN input variables show that the NN without convective memory as input suffers from numerical instability, and the NN without considering radiative variables and land fraction as input, or with reduced training samples produce less accurate results. To our knowledge, this is the first time an ML parameterization successfully achieves online extrapolation to a warm climate without using additional warm‐climate data for training. It demonstrates the potential of ML‐driven parameterizations for credible long‐term climate projections.

Atmosphere model↗

Perspectives on Systematic Cloud Microphysics Scheme Development With Machine Learning

Cloud microphysics—the collection of processes that govern the small‐scale formation, evolution, and interactions of liquid droplets and ice crystals in clouds and precipitation—remains a major source of uncertainty in weather and climate models. Although too small in scale to be explicitly resolved in any large‐eddy simulation, weather, or climate model, the representation of cloud microphysical processes has significant impact at the climate scale. Current microphysical schemes are limited by both parametric uncertainty, linked to uncertainty in physical parameter values, and structural uncertainty, arising from incomplete physical understanding of the processes at play or approximations made for computational efficiency. Recent advances in the application of machine learning (ML) to the physical sciences show significant potential for minimizing these limitations by leveraging high‐fidelity simulations and observations. Here we outline the challenges that must be addressed to apply ML toward cloud microphysics scheme development. This perspectives paper synthesizes recent progress in using data‐driven methods, including ML, to improve cloud microphysics parameterizations and highlights opportunities to address key uncertainties. We discuss the roles of aleatoric (irreducible, or statistical) and epistemic (reducible, or systematic) errors in contributing to microphysics parameterization uncertainty. ML can leverage observations to improve microphysical schemes via bottom‐up and top‐down constraints. Methods such as differentiable programming and ML‐enhanced sampling strategies and the creation of large scale benchmark data sets promise to bridge the gap between observations and models and to improve the consistency of cloud microphysical representation across temporal and spatial scales.

Lamb, Kara D. [Columbia Univ., New York, NY (Unite↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

Hydrology in the Age of Artificial Intelligence: From Fragmentation to Coherent Terrestrial Hydrosphere Science

The rapid rise of machine learning (ML) in hydrology has prompted debate about the discipline's scientific relevance. While ML often outperforms traditional models in streamflow prediction, we argue that this reflects a deeper limitation: persistent fragmentation of hydrological science itself. Narrow focus on isolated components has hindered the development of coherent, scale‐relevant understanding of the integrated terrestrial hydrosphere. This is illustrated, for example, by widely divergent estimates of groundwater–streamflow interactions and of water balance‐implied ongoing storage changes. We argue that hydrology's future lies not in choosing between ML and physics, but in integrating data‐driven and process‐based approaches to advance consistent, realistic, and societally relevant understanding of the terrestrial hydrosphere and its multifaceted roles in the Earth System.

Painter, Scott L. [Oak Ridge National Laboratory (↗

Analytical ab initio hessian from a deep learning potential for transition state optimization

Identifying transition states—saddle points on the potential energy surface connecting reactant and product minima—is central to predicting kinetic barriers and understanding chemical reaction mechanisms. In this work, we train a fully differentiable equivariant neural network potential, NewtonNet, on thousands of organic reactions and derive the analytical Hessians. By reducing the computational cost by several orders of magnitude relative to the density functional theory (DFT) ab initio source, we can afford to use the learned Hessians at every step for the saddle point optimizations. We show that the full machine learned (ML) Hessian robustly finds the transition states of 240 unseen organic reactions, even when the quality of the initial guess structures are degraded, while reducing the number of optimization steps to convergence by 2–3× compared to the quasi-Newton DFT and ML methods. All data generation, NewtonNet model, and ML transition state finding methods are available in an automated workflow.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Model-free estimation of completeness, uncertainties, and outliers in atomistic machine learning using information theory

Abstract An accurate description of information is relevant for a range of problems in atomistic machine learning (ML), such as crafting training sets, performing uncertainty quantification (UQ), or extracting physical insights from large datasets. However, atomistic ML often relies on unsupervised learning or model predictions to analyze information contents from simulation or training data. Here, we introduce a theoretical framework that provides a rigorous, model-free tool to quantify information contents in atomistic simulations. We demonstrate that the information entropy of a distribution of atom-centered environments explains known heuristics in ML potential developments, from training set sizes to dataset optimality. Using this tool, we propose a model-free UQ method that reliably predicts epistemic uncertainty and detects out-of-distribution samples, including rare events in systems such as nucleation. This method provides a general tool for data-driven atomistic modeling and combines efforts in ML, simulations, and physical explainability.

36 MATERIALS SCIENCE↗

Optical neural engine for solving scientific partial differential equations

Abstract Solving partial differential equations (PDEs) is the cornerstone of scientific research and development. Data-driven machine learning (ML) approaches are emerging to accelerate time-consuming and computation-intensive numerical simulations of PDEs. Although optical systems offer high-throughput and energy-efficient ML hardware, their demonstration for solving PDEs is limited. Here, we present an optical neural engine (ONE) architecture combining diffractive optical neural networks for Fourier space processing and optical crossbar structures for real space processing to solve time-dependent and time-independent PDEs in diverse disciplines, including Darcy flow equation, the magnetostatic Poisson’s equation in demagnetization, the Navier-Stokes equation in incompressible fluid, Maxwell’s equations in nanophotonic metasurfaces, and coupled PDEs in a multiphysics system. We numerically and experimentally demonstrate the capability of the ONE architecture, which not only leverages the advantages of high-performance dual-space processing for outperforming traditional PDE solvers and being comparable with state-of-the-art ML models but also can be implemented using optical computing hardware with unique features of low-energy and highly parallel constant-time processing irrespective of model scales and real-time reconfigurability for tackling multiple tasks with the same architecture. The demonstrated architecture offers a versatile and powerful platform for large-scale scientific and engineering computations.

Tang, Yingheng (ORCID:0009000153622546)↗