Search NASA⌕ Search

SEARCH · Search NASA

Results for “Learning with errors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Machine learning and process-based modeling of spatiotemporal changes in active layer thickness across Alaska

Permafrost degradation poses a growing threat to infrastructure stability and ecosystem resilience in the rapidly warming Arctic. We investigated the spatiotemporal dynamics of active layer thickness (ALT) across Alaska by integrating field observations, environmental datasets, a physically based Stefan model, and machine learning (ML) techniques. Using weather projections from the Coupled Model Intercomparison Project Phase 6 under two Shared Socioeconomic Pathways (SSP 2-4.5 and SSP 5-8.5), we assessed ALT sensitivity to projected future weather conditions. The random forest (RF) model outperformed the Stefan approach in predicting ALT on the training dataset (R² = 0.84 vs. 0.53) but demonstrated lower generalizability on the test dataset (R² = 0.24 vs. 0.54). The root mean square error (RMSE) for the RF model for training and testing ranged from 14 to 22 cm, compared to 17 and 18 cm for the Stefan model. Variable importance analysis revealed that mean annual temperature and slope angle were the strongest predictors of ALT, accounting for 19% and 18% of the variance, respectively, followed by sediment transport index (14%) and stream power index (11%). Comparative analysis of baseline ALT predictions showed the Stefan model tended to project a thicker active layer (mean ± SD: 65 ± 16 cm), compared to the RF model (mean ± SD: 59 ± 8.8) cm). Both models indicated a latitudinal gradient in ALT, with shallower depths at higher latitudes. Projected ALT increases by 2100 were estimated at 3.3 ± 2.2 cm under SSP 2-4.5 and 5.9 ± 4.0 cm under SSP 5-8.5 for the ML model, whereas the Stefan model projected substantially larger increases of 13 ± 2.6 cm (SSP 2-4.5) and 28 ± 4.4 cm (SSP5-8.5). Spatial analysis showed the greatest ALT increases in northern Alaska, with relatively smaller changes in southern regions. These findings highlight the complex, multifactorial nature of ALT dynamics and the value of hybrid modeling approaches. As rising temperatures accelerate permafrost thaw, changes in ALT can disrupt ecosystems, damage infrastructures, and enhance the release of stored soil carbon, highlighting the urgent need for improved predictive capabilities to inform adaptation strategies in the Arctic.

Climate sciences↗

Denoising Seismograms in the Time Domain Using a Deep Learning Model

Deep learning has emerged as a transformative tool for enhancing the extraction of reliable information from seismograms, addressing the increasing demand for precise and efficient seismic data analysis. We introduce an innovative encoder–decoder deep learning model, named WaveDenoiser, designed for noise reduction in the time domain, thereby eliminating the need for spectrogram computations that have been used for existing deep learning tools and significantly improving processing speed. Utilizing the benchmark dataset that is Stanford Earthquake Dataset, we developed three models of varying sizes: base, medium, and large. Notably, the large (referred to as WaveDenoiser) model demonstrated superior performance, achieving a median signal‐to‐noise ratio improvement of 8.8 dB on in‐distribution unseen data (in the same geographic region) and 7.7 dB on out‐distribution unseen data (in a new geographic region), outpacing both the base and medium models. Further evaluation of the WaveDenoiser model revealed a reduction in median arrival‐time errors by 0.02 s for P waves and 0.01 s for S waves when processing waveforms prior to phase picking using PhaseNet on in‐distribution unseen data. When tested on out‐distribution unseen data, the model also effectively reduced the P‐wave median arrival‐time error by 0.02 and 0.01 s in median arrival‐time error for S waves. Importantly, the application of WaveDenoiser resulted in a significant reduction of phase picking outliers by 1.1% to 3.6% for both P and S waves. In addition, we achieved over five times acceleration in processing speed compared with the seisBench implementation of DeepDenoiser. Our findings underscore the potential of WaveDenoiser as a powerful tool for improving seismic data analysis and processing efficiency.

P-waves↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

Virtual sensing-enabled digital twin framework for real-time monitoring of nuclear systems leveraging deep neural operators

Abstract Real-time monitoring is a foundation of nuclear digital twin technology, crucial for detecting material degradation and maintaining nuclear system integrity. Traditional physical sensor systems face limitations, particularly in measuring critical parameters in hard-to-reach or harsh environments, often resulting in incomplete data coverage. Machine learning-driven virtual sensors offer a transformative solution by complementing physical sensors in monitoring critical degradation indicators. This paper introduces the use of Deep Operator Networks (DeepONet) to predict key thermal-hydraulic parameters in the hot leg of pressurized water reactor. DeepONet acts as a virtual sensor, mapping operational inputs to spatially distributed system behaviors without requiring frequent retraining. Our results show that DeepONet achieves low mean squared and Relative L2 error, making predictions 1400 times faster than traditional CFD simulations . These characteristics enable DeepONet to function as a real-time virtual sensor, synchronizing with the physical system to track degradation conditions and provide insights within the digital twin framework for nuclear systems.

Hossain, Raisa↗

Signal-preserving CMB component separation with machine learning

Analysis of microwave sky signals, such as the cosmic microwave background, often requires component separation using multifrequency methods, whereby different signals are isolated according to their different frequency behaviors. Many so-called blind methods, such as the internal linear combination (ILC), make minimal assumptions about the spatial distribution of the signal or contaminants, and only assume knowledge of the frequency dependence of the signal. The ILC produces a minimum-variance linear combination of the measured frequency maps. In the case of Gaussian, statistically isotropic fields, this is the optimal linear combination, as the variance is the only statistic of interest. However, in many cases the signal we wish to isolate, or the foregrounds we wish to remove, are non-Gaussian and/or statistically anisotropic (in particular for the case of Galactic foregrounds). In such cases, it is possible that machine learning (ML) techniques can be used to exploit the non-Gaussian features of the foregrounds and thereby improve component separation. However, many ML techniques require the use of complex, difficult-to-interpret operations on the data. We propose a hybrid method whereby we train an ML model using only combinations of the data that , and combine the resulting ML-predicted foreground estimate with the ILC solution to reduce the error from the ILC. We demonstrate our methods on simulations of extragalactic temperature and Galactic polarization foregrounds and show that our ML model can exploit non-Gaussian features, such as point sources and spatially varying spectral indices, to produce lower-variance maps than ILC—e.g., reducing the variance of the B-mode residual by factors of up to 5—while preserving the signal of interest in an unbiased manner. Moreover, we often find improved performance even when applying our ML technique to foreground models on which it was not trained. Published by the American Physical Society 2025

McCarthy, Fiona (ORCID:0000000253893565)↗

Integration of the Biot–Gassmann Fluid Substitution Method and Machine Learning-Based Velocity–Stress Relationship for Estimating In Situ Stresses

Recent advancements have shown that in situ stresses can be reliably estimated through an integrated machine/deep learning (ML/DL)-based framework, which relies on models trained and validated using true triaxial ultrasonic velocity (TUV) experimental data that involve measurements of ultrasonic velocity in saturated rocks under varying stress configurations. However, when the goal is to interpret lower frequency measurements, it may be more appropriate to run experiments on dry rocks and then obtain Biot–Gassmann-derived equivalent saturated velocities (low-frequency approximation) and employ these quantities for training ML/DL models to predict in situ stress. Whether the dispersion effect of frequency on the velocity–stress relationship substantially impacts in situ stress prediction is an important and unresolved question. This work presents an enhancement of ML/DL-based workflow by training and implementing ML/DL models using equivalent saturated acoustic velocities (low-frequency) obtained by applying Biot–Gassmann fluid substitution on the ultrasonic velocities of dry cores. The models were trained on TUV data sets derived from three subsurface cores extracted from the geothermal well 16B(78)-32 at the Utah FORGE site. Each core was subjected to 75 unique stress configurations for velocity measurement in the dry state. The ML/DL trained on the TUV data set with equivalent saturated velocities demonstrated promising performance to predict in situ stress in subsurface geological rocks using velocity–stress relationships with R 2 of 0.86, 0.971, and 0.975 and root mean squared error (RMSE) of 2.59, 1.92, and 1.80 for validation/testing phases of vertical, minimum horizontal, and maximum horizontal stress models, respectively. Additionally, interpretation and explanation by Shapley additive explanations (SHAP) analysis further improved scientific validation and model reliability for estimating in situ stresses.

colloids↗

High-Fidelity Accelerated Design of High-performance Electrochemical Systems

Large-scale electrification is vital to addressing the climate crisis, but several scientific and technological challenges remain to fully electrify both the chemical industry and transportation. In both of these areas, new electrochemical materials will be critical, but their development currently relies heavily on human-time-intensive experimental trial and error and computationally expensive first-principles, meso-scale and continuum simulations. To accelerate this process, our team has developed the AutoMat platform. AutoMat can accelerate development of new electrochemical materials along two avenues: first, automated input generation and management of simulations at multiple lengthscales as well as “handoff” of outputs from one lengthscale as inputs to the next; and second, replacement of the most computationally intensive simulation processes with machine-learned surrogate models. The crux of our team’s effort was not “reinventing the wheel” by developing entirely new techniques, but rather building a “superhighway” that allows existing state-of-the-art techniques to run faster and more smoothly than before. AutoMat can utilize tools spanning from first-principles quantum chemistry computations to automated robotic experimentation, and is driven by design space search techniques to reduce the number of iterations through the full simulation loop by rapidly targeting promising regions of design spaces such as single-atom alloy catalysts or blends of liquid electrolytes.

25 ENERGY STORAGE↗

Machine Learning Techniques for Data Reduction of Climate Applications

Scientists conduct large-scale simulations to compute derived quantities-of-interest (QoI) from primary data. Often, QoI are linked to specific features, regions, or time intervals, such that data can be adaptively reduced without compromising the integrity of QoI. For many spatiotemporal applications, these QoI are binary in nature and represent presence or absence of a physical phenomenon. We present a pipelined compression approach that first uses neural-network-based techniques to derive regions where QoI are highly likely to be present. Then, we employ a Guaranteed Autoencoder (GAE) to compress data with differential error bounds. GAE uses QoI information to apply low-error compression to only these regions. This results in overall high compression ratios while still achieving downstream goals of simulation or data collections. Experimental results are presented for climate data generated from the E3SM Simulation model for downstream quantities such as tropical cyclone and atmospheric river detection and tracking. These results show that our approach is superior to comparable methods in the literature.

Li, Xiao [University of Florida]↗

A machine learning approach to quantify degradation of nuclear fuels and the effects of fission products

Nuclear fuel performance is critically dependent on understanding the evolution of fuel properties under operational conditions, a complex challenge driven by chemical changes and substantial radiation damage during fission. Traditionally, property evolution has been determined via empirical data collected following irradiation. However, these empirical correlations are limited in their applicability beyond the specific conditions in which they were obtained. This study explores a novel approach to address this challenge by applying materials informatics to develop a machine learning random forest (ML-RF) model that captures the effects of fission products on fuel compounds. The model predicts formation enthalpy (ΔH f ) by leveraging extensive quantum materials property data and correlating it with material descriptors such as composition, atomic and site features, and crystal lattice properties. This ML-RF model enables rapid interpolation across the compositional and structural spaces covered by the training data, thus supporting high-throughput screening and energetic ranking of candidate phases. The model demonstrates the ability to predict ΔH f with a mean absolute error (MAE) of approximately 0.1 to 0.2 eV/atom across a wide range of compounds, including key nuclear fuel systems (U-O, U-N, U-C, U-Si, and U-Mo). For example, it was used to assess shifts in stoichiometry for UO 2 (O/M) and UN (N/M) fuels, revealing their distinct tendencies in chemical potential variation and enabling preliminary convex hull analyses. Furthermore, the model provides insights into how individual fission products affect fuel properties. Results indicate that larger fission products (e.g., Nd, Pu, Ce) have a more pronounced impact on UO 2 , while lighter ones (e.g., Zr) strongly influence UN. Here, the model developed in this work can be used to support the Accelerated Fuel Qualification approach by facilitating preliminary evaluations prior to extensive materials modeling and experimentation. To this end, the trained model has been made available to the fuel community to support ongoing fuel development efforts.

Accelerated fuel qualification↗

Multi-modality deep learning for pulse prediction in homogeneous nonlinear systems via parametric conversion

In this Letter, we introduce FusionNet, a multi-modality deep learning framework designed to predict and analyze output pulses in high-power rare-earth-doped laser systems driving parametric conversion in homogeneous guided nonlinear media. FusionNet integrates temporal, spectral, and physical experimental conditions to model ultrafast nonlinear phenomena, including parametric nonlinear frequency conversion, self-phase modulation, and cross-phase modulation in homogeneous guided systems such as gas-filled hollow-core fibers. These systems bridge physical models with experimental data, advancing our understanding of light-guiding principles and nonlinear interactions while expediting the design and optimization of on-demand high-power, high-brightness systems. Our results demonstrate a 73% reduction in prediction error and an 83% improvement in computational efficiency compared to conventional neural networks. This work establishes a new paradigm for accelerating parametric simulations and optimizing experimental designs in high-power laser systems, with further implications for high-precision spectroscopy, quantum information science, and distributed entangled interconnects.

47 OTHER INSTRUMENTATION↗

Prediction of the Cu oxidation state from EELS and XAS spectra using supervised machine learning

Abstract Electron energy loss spectroscopy (EELS) and X-ray absorption spectroscopy (XAS) provide detailed information about bonding, distributions and locations of atoms, and their coordination numbers and oxidation states. However, analysis of XAS/EELS data often relies on matching an unknown experimental sample to a series of simulated or experimental standard samples. This limits analysis throughput and the ability to extract quantitative information from a sample. In this work, we have trained a random forest model capable of predicting the oxidation state of copper based on its L-edge spectrum. Our model attains an R 2 score of 0.85 and a root mean square error of 0.24 on simulated data. It has also successfully predicted experimental L-edge EELS spectra taken in this work and XAS spectra extracted from the literature. We further demonstrate the utility of this model by predicting simulated and experimental spectra of mixed valence samples generated by this work. This model can be integrated into a real-time EELS/XAS analysis pipeline on mixtures of copper-containing materials of unknown composition and oxidation state. By expanding the training data, this methodology can be extended to data-driven spectral analysis of a broad range of materials.

36 MATERIALS SCIENCE↗

Next generation Arctic vegetation maps: Aboveground plant biomass and woody dominance mapped at 30 m resolution across the tundra biome

The Arctic is warming faster than anywhere else on Earth, placing tundra ecosystems at the forefront of global climate change. Plant biomass is a fundamental ecosystem attribute that is sensitive to changes in climate, closely tied to ecological function, and crucial for constraining ecosystem carbon dynamics. However, the amount, functional composition, and distribution of plant biomass are only coarsely quantified across the Arctic. Therefore, we developed the first moderate resolution (30 m) maps of live aboveground plant biomass (g m −2 ) and woody plant dominance (%) for the Arctic tundra biome, including the mountainous Oro Arctic. We modeled biomass for the year 2020 using a new synthesis dataset of field biomass harvest measurements, Landsat satellite seasonal synthetic composites, ancillary geospatial data, and machine learning models. Additionally, we quantified pixel-wise uncertainty in biomass predictions using Monte Carlo simulations and validated the models using a robust, spatially blocked and nested cross-validation procedure. Observed plant and woody plant biomass values ranged from 0 to ∼6000 g m −2 (mean ≈ 350 g m −2 ), while predicted values ranged from 0 to ∼4000 g m −2 (mean ≈ 275 g m −2 ), resulting in model validation root-mean-squared-error (RMSE) ≈ 400 g m −2 and R 2 ≈ 0.6. Our maps not only capture large-scale patterns of plant biomass and woody plant dominance across the Arctic that are linked to climatic variation (e.g., thawing degree days), but also illustrate how fine-scale patterns are shaped by local surface hydrology, topography, and past disturbance. By providing data on plant biomass across Arctic tundra ecosystems at the highest resolution to date, our maps can significantly advance research and inform decision-making on topics ranging from Arctic vegetation monitoring and wildlife conservation to carbon accounting and land surface modeling.

Climate change↗

A machine learning method of modern urban building energy modeling: A case study of Chicago

Urban-scale building energy modeling is vital for urban planning. However, it can be challenging to assimilate reliable non-geometry building data for urban-scale modeling without extensive investment. Here, this study introduces a novel approach to developing modern urban-scale building energy stock data using geographic information systems and machine learning algorithms without necessarily requiring pre-supplied non-geometric metadata. The proposed framework integrates building footprint and height data to estimate gross floor areas, and matches each building to a pool of candidate records from ComStock or ResStock—filtered to the same county and ranked by geometric similarity—demonstrate a proof-of-concept case study in Chicago for predicting energy use intensity (EUI) using scalable datasets. The model achieved a mean bias error (MBE) of 0.08 kWh/m² and root mean square error (RMSE) of 14.84 kWh/m² under full metadata input for EUI prediction. With only location inputs, the model captured 69.2 % of EUI within predicted ranges. These results demonstrate the model’s potential to support early-stage urban planning, identify candidates for energy-efficient retrofits. By removing the dependency on detailed pre-surveys or extensive building metadata, the approach overcomes a key barrier in traditional urban-scale building energy modeling, illustrating a pathway toward broader and more cost-effective application, though further multi-city validation and improved treatment of pre-1925 buildings are needed.

Energy Use Intensity↗

Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO

Abstract Streambed grain sizes control river hydro‐biogeochemical (HBGC) processes and functions. However, measuring their quantities, distributions, and uncertainties is challenging due to the diversity and heterogeneity of natural streams. This work presents a photo‐driven, artificial intelligence (AI)‐enabled, and theory‐based workflow for extracting the quantities, distributions, and uncertainties of streambed grain sizes from photos. Specifically, we first trained You Only Look Once, an object detection AI, using 11,977 grain labels from 36 photos collected from nine different stream environments. We demonstrated its accuracy with a coefficient of determination of 0.98, a Nash–Sutcliffe efficiency of 0.98, and a mean absolute relative error of 6.65% in predicting the median grain size of 20 ground‐truth photos representing nine typical stream environments. The AI is then used to extract the grain size distributions and determine their characteristic grain sizes, including the 10th, 50th, 60th, and 84th percentiles, for 1,999 photos taken at 66 sites within a watershed in the Northwest US. The results indicate that the 10th, median, 60th, and 84th percentiles of the grain sizes follow log‐normal distributions, with most likely values of 2.49, 6.62, 7.68, and 10.78 cm, respectively. The average uncertainties associated with these values are 9.70%, 7.33%, 9.27%, and 11.11%, respectively. These data allow for the computation of the quantities, distributions, and uncertainties of streambed HBGC parameters, including Manning's coefficient, Darcy‐Weisbach friction factor, top layer interstitial velocity magnitude, and nitrate uptake velocity. Additionally, major sources of uncertainty in grain sizes and their impact on HBGC parameters are examined.

58 GEOSCIENCES↗

Intercomparison of Deep Learning Model Architectures for Atmospheric River Prediction

With a rapid surge in the application of machine learning (ML) for a diverse range of tasks in climate science, the present study addresses a challenge for climate scientists when selecting the optimal ML or deep learning (DL) architecture for a given application. In particular, a DL intercomparison study was performed with a focus on forecasting the position of atmospheric rivers (ARs) on short-range time scales (up to 5-day lead times). AR predictions from multiple DL architectures, including various types of convolutional autoencoders and a vision transformer (ViT), were compared against ECMWF ERA5 reanalysis and hindcasts from a global climate model. DL models with similar trainable parameters were trained on ERA5 reanalysis data and AR positions derived from a thresholding algorithm to ensure a fair comparison among the DL models. Each model’s performance and accuracy in forecasting AR location and key input fields within a 5-day window were assessed using metrics of root-mean-square error, anomaly correlation, and mean intersection over union. The ViT architecture outperformed other autoencoder models in most of the metrics. Incorporating additional meteorological fields only yielded slight improvements in forecasting certain fields at longer lead times. The results also suggest that a smaller number of input time steps or smaller number of autoregressive steps can achieve better prediction skills, while also improving the overall computational efficiency. This research offers valuable insights into the strengths and weaknesses of different DL techniques for AR forecasting, hopefully guiding the development of improved models for forecasting this phenomenon.

54 ENVIRONMENTAL SCIENCES↗

Deep Learning Reconstruction of Daily Soil CO 2 Efflux Reveals Biogeochemical Insights and Reduces Annual Estimate Uncertainty Despite Limited Daily Predictability

Soil CO 2 efflux is commonly measured monthly or seasonally, leaving daily dynamics poorly resolved and contributing to global estimation uncertainty. We trained a single Long Short-Term Memory (LSTM) model to predict daily soil CO 2 efflux across 82 globally distributed sites in COSORE, with 0.2%–46.9% daily data coverage from 2003 to 2020. Despite using far fewer sites than are typically used to train a single deep learning model, with observations biased toward temperate mesic sites, the LSTM model performed well at approximately one-third of sites, reconstructed nearly 2 decades of daily efflux, and outperformed commonly used approaches for estimating daily efflux when applied to the same data set. Performance was weakest at pronounced peaks and troughs and at non-temperate sites with <1.5 years of observations and irregular data patterns. Nevertheless, annual efflux from reconstructed daily data had <40% error even at underperforming sites, substantially improving estimates derived from monthly and seasonal sampling (maximum errors of 95% and 136%, respectively). Temperature sensitivity (Q 10 ) estimated from reconstructed daily predictions closely matched estimates from daily observations, whereas Q 10 values derived from monthly or seasonal observations deviated substantially, suggesting that coarse temporal sampling may contribute to uncertainty in reported Q 10 values. Consistent daily reconstructions further enabled trend analyses for well-performing, predominantly temperate sites and showed increasing soil CO 2 efflux at most sites from 2003 to 2020, with more variable summer trends. Despite limitations, these results demonstrate the potential of LSTM models to reconstruct daily soil CO 2 efflux and reduce estimation uncertainties from sparse observations.

Smykalov, Valerie [Pennsylvania State University, ↗

Perspectives on Systematic Cloud Microphysics Scheme Development With Machine Learning

Cloud microphysics—the collection of processes that govern the small‐scale formation, evolution, and interactions of liquid droplets and ice crystals in clouds and precipitation—remains a major source of uncertainty in weather and climate models. Although too small in scale to be explicitly resolved in any large‐eddy simulation, weather, or climate model, the representation of cloud microphysical processes has significant impact at the climate scale. Current microphysical schemes are limited by both parametric uncertainty, linked to uncertainty in physical parameter values, and structural uncertainty, arising from incomplete physical understanding of the processes at play or approximations made for computational efficiency. Recent advances in the application of machine learning (ML) to the physical sciences show significant potential for minimizing these limitations by leveraging high‐fidelity simulations and observations. Here we outline the challenges that must be addressed to apply ML toward cloud microphysics scheme development. This perspectives paper synthesizes recent progress in using data‐driven methods, including ML, to improve cloud microphysics parameterizations and highlights opportunities to address key uncertainties. We discuss the roles of aleatoric (irreducible, or statistical) and epistemic (reducible, or systematic) errors in contributing to microphysics parameterization uncertainty. ML can leverage observations to improve microphysical schemes via bottom‐up and top‐down constraints. Methods such as differentiable programming and ML‐enhanced sampling strategies and the creation of large scale benchmark data sets promise to bridge the gap between observations and models and to improve the consistency of cloud microphysical representation across temporal and spatial scales.

Lamb, Kara D. [Columbia Univ., New York, NY (Unite↗

Machine Learning-Enabled Image Classification for Automated Electron Microscopy

Abstract Traditionally, materials discovery has been driven more by evidence and intuition than by systematic design. However, the advent of “big data” and an exponential increase in computational power have reshaped the landscape. Today, we use simulations, artificial intelligence (AI), and machine learning (ML) to predict materials characteristics, which dramatically accelerates the discovery of novel materials. For instance, combinatorial megalibraries, where millions of distinct nanoparticles are created on a single chip, have spurred the need for automated characterization tools. This paper presents an ML model specifically developed to perform real-time binary classification of grayscale high-angle annular dark-field images of nanoparticles sourced from these megalibraries. Given the high costs associated with downstream processing errors, a primary requirement for our model was to minimize false positives while maintaining efficacy on unseen images. We elaborate on the computational challenges and our solutions, including managing memory constraints, optimizing training time, and utilizing Neural Architecture Search tools. The final model outperformed our expectations, achieving over 95% precision and a weighted F-score of more than 90% on our test data set. This paper discusses the development, challenges, and successful outcomes of this significant advancement in the application of AI and ML to materials discovery.

Materials Science↗