Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine learning models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Advancing stream temperature prediction with a generalizable large-sample framework across CONUS river reaches

Accurately predicting stream temperature in ungauged basins remains a critical challenge for water resource management, thermoelectric power plant cooling, and ecosystem conservation. Large-sample machine learning models trained on hundreds of well-monitored river basins have shown remarkable performance; however, such models have yet to be developed solely using forcing data that can be readily extracted to simulate stream temperatures anywhere in the contiguous United States (CONUS). In this study, we present a scalable, large-sample deep learning framework using Long Short-Term Memory (LSTM) networks to simulate daily stream temperatures in ungauged basins across the CONUS. The framework leverages both modeled reanalysis of meteorological and streamflow inputs as well as static attributes available for all 2.7 million CONUS river reaches in the National Hydrography Dataset Plus (NHDPlusV2). By generating dynamical inputs from predefined thermally relevant upstream contributing areas, rather than the entire upstream basin, the model also offers improvements in very large basins where full-basin averaging can dilute the most important influences on stream temperature. Evaluated across 300 basins, the model achieves a median Mean Absolute Error (MAE) of 1.1 °C and a Nash-Sutcliffe Efficiency (NSE) of 0.95 on temporally and spatially distinct test folds—comparable to models trained exclusively using meteorological and streamflow observational data. The flexible, high-performing framework generalizes to any unmonitored river reach without significant regulation or unnatural thermal input immediately upstream, substantially expanding predictive capabilities in data-scarce regions.

Hydrology↗

Sensor Reduction for Diversion Detection in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors are designed as a smaller, cheaper, and safer alternative to traditional nuclear power plants. Their non-traditional characteristics and prospect of mass production and deployment will likely require new approaches to nuclear safeguards. The primary proliferation concern with microreactors is the diversion of fuel material. Such diversion may produce measurable defects in key physical attributes like neutron flux, which may in turn be detectable using machine learning models. Preliminary work has demonstrated this ability for modeled nominal and diversion scenarios using large quantities of energy integrated neutron flux data. In practice, the number of available sensors for such measurements will be limited and energy integrated flux information will not be available. This work explores the ability of tree-based gradient boosted ensemble models to classify a given microreactor core is nominal or diversion, and determine the number of fuel pins diverted in the case of diversion with reduced numbers of sensors and more realistic detector responses. Classification accuracy of greater than 98% and regression errors as low as 5% of the total number of fuel pins were achieved with as few as 15 sensors, compared to 99% and 4.1% with a maximum of 240 sensors.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Investigating the opioid epidemic across the United States: Associations between county-level characteristics and overdose mortality

The opioid crisis remains a critical public health challenge in the United States. Despite national efforts that reduced opioid prescribing by nearly 44% between 2011 and 2021, opioid overdose deaths more than tripled during the same period. This alarming trend reflects a major shift in the crisis, with illegal opioids now driving the majority of overdose deaths instead of prescription opioids. Although supply-side factors fueling this transition have been widely studied, the structural and community-level conditions that shape overdose mortality are less well understood. To help address this gap, this study has three primary objectives: (1) overcome structural gaps in national data to construct a complete nationwide county-level dataset from 2010 to 2022; (2) using data analysis, identify and investigate spatiotemporal anomalies in overdose mortality; and (3) using two machine-learning models, quantify the importance of thirteen social vulnerability variables in predicting overdose mortality. Our results identify unemployment and limited vehicle access as key county-level predictors of overdose mortality. Higher levels of these vulnerabilities are associated with elevated mortality, whereas lower levels are associated with reduced mortality. These findings highlight factors that may be relevant for public health planning and policy prioritization within the context of the opioid crisis.

Anomaly analysis↗

Fundamental limit of jet tagging

Identifying the origin of high-energy hadronic jets (jet tagging) has been a critical benchmark problem for machine learning in particle physics. Jets are ubiquitous at colliders and are complex objects that serve as prototypical examples of collections of particles to be categorized. Over the last decade, machine learning-based classifiers have replaced classical observables as the state of the art in jet tagging. Increasingly complex machine learning models are leading to increasingly more effective tagger performance. Our goal is to address the question of convergence—are we getting close to the fundamental limit on jet tagging or is there still potential for computational, statistical, and physical insights for further improvements? We address this question using state-of-the-art generative models to create a realistic, synthetic dataset with a known jet tagging optimum. Various state-of-the-art taggers are deployed on this dataset, showing that there is a significant gap between their performance and the optimum. Our dataset and software are made public to provide a benchmark task for future developments in jet tagging and other areas of particle physics.

Artificial intelligence↗

Classification of Cloud Particle Imagery and Thermodynamics (COCPIT): A New Databasing Tool for the Characterization of Cloud Particle Images Captured During DOE Field Campaigns

The Department of Energy for decades has explored the earth system and atmosphere through research and deployment of in-situ and remote sensing platforms during field campaigns. Among these datasets exists a vast supply of cloud particle images that provide visual insight into the complex microphysics in the clouds that span our globe. The millions of images collected over decades of deployments provides a unique opportunity to further our understanding of our atmosphere down to the crystal size. This work over the past 5 years has sought to organize these images into digestible datasets that can then be used by scientists to further our understanding of microphysics. A machine learning model was developed that categorizes over 1.5 million images across 11 weather events with over 90% accuracy according to particle type. The database was then extended to include dimensional characteristics of the particle as well as co-location of environmental properties, such as temperature and water content. Then, to initialize the connection between these data and our understanding of how crystals form and grow, weather research and forecasting simulations were run to generate the growth histories of the classified crystals. This research culminates with 2 databases per event: (1) a database of all classified crystals and their dimensional and environmental properties and (2) simulated growth histories of each crystal. Finally, a user interface was created to allow researchers to explore data statistics.

54 ENVIRONMENTAL SCIENCES↗

Establishing nationwide power system vulnerability index across US counties using interpretable machine learning

Power outages have become increasingly frequent, intense, and prolonged in the US due to climate change, aging electrical grids, and rising energy demand. However, largely due to the absence of granular spatiotemporal outage data, we lack data-driven evidence and analytics-based metrics to quantify power system vulnerability. This limitation has hindered the ability to effectively evaluate and address vulnerability to power outages in US communities. Here, in this work, we collected ∼179 million power outage records at 15-min intervals across 3022 US contiguous counties (96.15 % of the area) from 2014 to 2023. We developed a power system vulnerability assessment framework based on three dimensions (intensity, frequency, and duration) and applied interpretable machine learning models (XGBoost and SHAP) to compute Power System Vulnerability Index (PSVI) at the county level. Our analysis reveals a consistent increase in power system vulnerability across the US counties over the past decade. We identified 318 counties across 45 states as hotspots for high power system vulnerability, particularly in the West Coast (California and Washington), the East Coast (Florida and the Northeast area), the Great Lakes megalopolis (Chicago-Detroit metropolitan areas), and the Gulf of Mexico (Texas). Our heterogeneity analysis indicates that urban counties and those located along regional transmission boundaries tend to exhibit significantly higher vulnerability. Our results highlight the significance of the proposed PSVI for evaluating the vulnerability of communities to power outages. The findings underscore the widespread and pervasive impact of power outages across the country and offer crucial insights to support infrastructure operators, policymakers, and emergency managers in formulating policies and programs aimed at enhancing the resilience of the US power infrastructure.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Critical statistical assessment of data in metal additive manufacturing

Obtaining high quality data reflecting the relationships between the additive manufacturing (AM) process parameters, material microstructure and mechanical properties is crucial for the use of machine learning in AM. A database of over 4,000 data entries of metal AM was created thanks to a large number of literature studies on key process parameters and indicators of build quality. Meta-analysis reveals critical biases in the literature. Firstly, majority of studies report only high quality builds, these imbalances in reporting result in weak correlation between process parameters, properties and consolidation, limiting the ability of machine learning models to generalize beyond optimized conditions. Nevertheless, the trained models accurately predict yield strength ($R^2 = 0.85$), suggesting that certain process–property relationships are effectively captured within these models. Secondly, quantitative microstructural data are largely absent, limiting the learning of the microstructure-mechanical properties relationships. Finally, current process window identification is based largely on the consolidation, despite significant uncertainty in its measurement. It is important to identify the process map on the basis of not only the consolidation, but also mechanical behaviour under loading. Such a identification shows that 316 L and Inconel have much larger process map (i.e. highly printable) in comparison to the AlSi10Mg and Ti6Al4V.

Additive manufacturing↗

Quantifying the basic reproduction number and underestimated fraction of Mpox cases worldwide at the onset of the outbreak

In 2022, there was a global resurgence of mpox, with different clinical-epidemiological features compared with previous outbreaks. Sexual contact was hypothesized as the primary transmission route, and the community of men having sex with men (MSM) was disproportionately affected. Because of the stigma associated with sexually transmitted infections, the real burden of mpox could be masked. We quantified the basic reproduction number (R 0 ) and the underestimated fraction of mpox cases in 16 countries, from the onset of the outbreak until early September 2022, using Bayesian inference and a compartmentalized, risk-structured (high-/low-risk populations) and two-route (sexual/non-sexual transmission) mathematical model. Machine learning (ML) was harnessed to identify underestimation determinants. Estimated R 0 ranged between 1.37 (Canada) and 3.68 (Germany). The underestimation rates for the high- and low-risk populations varied between 25–93% and 65–85%, respectively. The estimated total number of mpox cases, relative to the reported cases, is highest in Colombia (3.60) and lowest in Canada (1.08). In the ML analysis, two clusters of countries could be identified, differing in terms of attitudes towards the 2SLGBTQIAP+ community and the importance of religion. Given the substantial mpox underestimation, surveillance should be enhanced, and country-specific campaigns against the stigmatization of MSM should be organized, leveraging community-based interventions.

60 APPLIED LIFE SCIENCES↗

Early calendar life and health prediction of silicon batteries via machine learning with uncertainty quantification

Lithium-ion batteries with silicon anodes promise high energy density but are limited by calendar lifetime. Reducing the long iteration time to obtain experimental results requires predicting calendar lifetime early in a cell's life. In this study, we demonstrate that lightweight machine learning models with feature engineering can provide calendar lifetime estimates from early electrochemical signals. After 1 month of electrochemical aging, the best models achieve 10% error in calendar-life prediction and can separate "bad" from "good" lifetime cells with a mean F1 score of 0.857. As battery systems exhibit inherent variability, four methods for uncertainty quantification are compared, and confidence intervals are demonstrated with an uncertainty of +-3.6 months in lifetime prediction. A feature importance analysis indicates that early patterns in voltage decay are the strongest indicators of calendar lifetime. Finally, this modeling approach has high error when generalizing to new electrode chemistries or testing conditions but with appropriately low confidence.

25 ENERGY STORAGE↗

Coincidence anomaly detection for unsupervised locating of edge localized modes in the DIII-D tokamak dataset

Using supervised learning to train a machine learning model to predict an on-coming edge localized mode (ELM) requires a large number of labeled samples. Creating an appropriate data set from the very large database of discharges at a long-running tokamak, such as DIII-D, would be a very time-consuming process for a human. Considering this need and difficulty, we use coincidence anomaly detection, an unsupervised learning technique, to train an ELM-identifier to identify and label ELMs in the DIII-D discharge database. This ELM-identifier shows, simultaneously, a precision of 0.68 and a recall of 0.63 (AUC is 0.73) on identifying ELMs in example time series pulled from thousands of discharges spanning five years. In a test set of 50 discharges, the algorithm finds over 26 thousand ELM candidates, more than 5 times the existing catalog of ELMs labeled by humans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Supporting ARPA-E Power Grid Optimization (Final Report)

Pacific Northwest National Laboratory (PNNL), Arizona State University (ASU), Georgia Institute of Technology (Georgia Tech), Los Alamos National Laboratory (LANL), National Renewable Energy Laboratory (NREL), Texas A&M University (TAMU), The University of Texas at Austin (UT), and the University of Wisconsin-Madison (UW-M) supported the ARPA-E Grid Optimization (GO) Competition by providing a common problem formulation, data format, datasets, evaluation mechanism, scoring, rules, and results that resulted in the awarding of $\$9.24$ million dollars to teams from academia, industry, and national labs for solving three sets of increasingly difficult non-linear, security- constrained AC Optimal Powerflow (AC-OPF) optimization problems in order to increase the efficiency of the US Electric Grid. It is estimated that a 1% increase in efficiency can save $\$1$ billion. Current industry practices typically use a linear DC model (DC-OPF) in order solve the OPF problem within the time constraints of the operation schedule. The GO Competition challenges the best power engineers, mathematicians, and computer scientists to make possible operational decisions based on accurate physical models. To accomplish this, the GO Competition created a series of Challenges and funded teams to produce the best solver. Challenge 1 was to solve the security constrained Alternating Current Optimal Power Flow (ACOPF) problem. Challenge 2 extended that to by adding adjustable transformer tap ratios, phase shifting transformers, switchable shunts, price-responsive demand, ramp rate constrained generators and loads, and fast-start unit commitment (UC). Furthermore, Challenge 2 was a maximization problem while Challenge 1 was a minimization problem. While Challenge 3 was being developed, the entrants were invited to find better solutions to the Challenge 2 synthetic datasets with no restrictions on time, hardware, or algorithms. The Challenge 2 solutions turned out to be very good. Challenge 3 expanded the Challenge 2 problem further by using multiperiod dynamic markets, including advisory models for extreme weather events, day-ahead markets, and the real-time markets with an extended look-ahead. These problems included active bid-in demand and topology optimization. Together the Challenges used nearly 30 million CPU hours. Since each team was working on the same problem, using the same data, and running on the same hardware, fair comparisons could be drawn as to the best solver. The datasets were varied enough, however, that the best solver for one dataset was not necessarily the best at another, so cumulative scores were used. The process was managed by the PNNL maintained website https://GOCompetition.energy.gov, where Entrants could find information about the problem, the data, the rules, submit their solver for evaluation, and see the scores of all the competing teams on a Leaderboard. Interest was world-wide but only American teams were eligible for prizes. The Competition has produced 34 journal articles 115 papers and been cited over 500 times in the literature, including 12 dissertations (4 from foreign countries; Columbia (2), Germany, and Italy) and 3 from the DOE ExaScale project. Software developed by Pearl Street Technologies for Challenges 1 and 2 is now deployed by Southwest Power Pool (SPP) and Midcontinent Independent Service Operator (MISO). Other teams have received inquiries from venture capitalists. Google DeepMind has thanked the Competition for making the datasets developed for the Competition public. They are using it to train machine learning models. The larger datasets have billions of unknowns to be solved for, but only a small percent matter in the final solution. Knowing what unknowns are important can dramatically speedup the solution.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep generative learning of magnetic frustration in artificial spin ice from magnetic force microscopy images

Increasingly large datasets of microscopic images with nanoscale resolution facilitate the development of machine learning methods to identify and analyze subtle physical phenomena embedded within the images. In this work, microscopic images of honeycomb lattice spin-ice samples serve as datasets from which we automate the calculation of net magnetic moments and directional orientations of spin-ice configurations. In the first stage of our workflow, machine learning models are trained to accurately predict magnetic moments and directions within spin-ice structures. Variational Autoencoders (VAEs), an emergent unsupervised deep learning technique, are employed to generate high-quality synthetic magnetic force microscopy (MFM) images and extract latent feature representations, thereby reducing experimental and segmentation errors. The second stage of proposed methodology enables precise identification and prediction of frustrated vertices and nanomagnetic segments, effectively correlating structural and functional aspects of microscopic images. This facilitates the design of optimized spin-ice configurations with controlled frustration patterns, enabling potential on-demand synthesis.

36 MATERIALS SCIENCE↗

Persistent Classification: Understanding Adversarial Attacks by Studying Decision Boundary Dynamics

ABSTRACT There are a number of hypotheses underlying the existence of adversarial examples for classification problems. These include the high‐dimensionality of the data, the high codimension in the ambient space of the data manifolds of interest, and that the structure of machine learning models may encourage classifiers to develop decision boundaries close to data points. This article proposes a new framework for studying adversarial examples that does not depend directly on the distance to the decision boundary. Similarly to the smoothed classifier literature, we define a (natural or adversarial) data point to be ( γ , σ)‐stable if the probability of the same classification is at least for points sampled in a Gaussian neighborhood of the point with a given standard deviation . We focus on studying the differences between persistence metrics along interpolants of natural and adversarial points. We show that adversarial examples have significantly lower persistence than natural examples for large neural networks in the context of the MNIST and ImageNet datasets. We connect this lack of persistence with decision boundary geometry by measuring angles of interpolants with respect to decision boundaries. Finally, we connect this approach with robustness by developing a manifold alignment gradient metric and demonstrating the increase in robustness that can be achieved when training with the addition of this metric.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bayesian Adaptive Polynomial Chaos Expansions

Polynomial chaos expansions (PCEs) are widely used for uncertainty quantification (UQ) tasks, particularly in the applied mathematics community. However, PCE has received comparatively less attention in the statistics literature, and fully Bayesian formulations remain rare—especially with implementations in R. Motivated by the success of adaptive Bayesian machine learning models such as BART, BASS and BPPR, we develop a new fully Bayesian adaptive PCE method with an efficient and accessible R implementation: khaos. Our approach includes a novel proposal distribution that enables data-driven interaction selection and supports a modified g-prior tailored to PCE structure. Through simulation studies and real-world UQ applications, we demonstrate that the Bayesian adaptive PCE provides competitive performance for surrogate modeling, global sensitivity analysis and ordinal regression tasks.

97 MATHEMATICS AND COMPUTING↗

Ensemble‐Based, Large‐Eddy Reconstruction of Wind Turbine Inflow in a Near‐Stationary Atmospheric Boundary Layer Through Generative Artificial Intelligence

ABSTRACT To validate the second‐by‐second dynamics of turbines in field experiments, it is necessary to accurately reconstruct the winds going into the turbine. Current time‐resolved inflow reconstruction techniques estimate wind behavior in unobserved regions using relatively simple spectral‐based models of the atmosphere. Here, we develop a technique for time‐resolved inflow reconstruction that is rooted in a large‐eddy simulation model of the atmosphere. Our “large‐eddy reconstruction” technique blends observations and atmospheric model information through a diffusion model machine learning algorithm, allowing us to generate probabilistic ensembles of reconstructions for a single 10‐min observational period. Our generated inflows can be used directly by aeroelastic codes or as inflow boundary conditions in a large‐eddy simulation. We verify the second‐by‐second reconstruction capability of our technique in three synthetic field campaigns, finding positive Pearson correlation coefficient values () between ground‐truth and reconstructed streamwise velocity, as well as smaller positive correlation coefficient values for unobserved fields (spanwise velocity, vertical velocity, and temperature). We validate our technique in three real‐world case studies by driving large‐eddy simulations with reconstructed inflows and comparing to independent inflow measurements. The reconstructions are visually similar to measurements, follow desired power spectra properties, and track second‐by‐second behavior ().

17 WIND ENERGY↗

Non-resonant anomaly detection with background extrapolation

Complete anomaly detection strategies that are both signal sensitive and compatible with background estimation have largely focused on resonant signals. Non-resonant new physics scenarios are relatively under-explored and may arise from off-shell effects or final states with significant missing energy. In this paper, we extend a class of weakly supervised anomaly detection strategies developed for resonant physics to the non-resonant case. Machine learning models are trained to reweight, generate, or morph the background, extrapolated from a control region. A classifier is then trained in a signal region to distinguish the estimated background from the data. The new methods are demonstrated using a semi-visible jet signature as a benchmark signal model, and are shown to automatically identify the anomalous events without specifying the signal ahead of time.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-throughput validation of phase formability and simulation accuracy of Cantor alloys

High-throughput methods enable accelerated discovery of novel materials in complex systems such as high-entropy alloys, which exhibit intricate phase stability across vast compositional spaces. Computational approaches, including Density Functional Theory (DFT) and calculation of phase diagrams (CALPHAD), facilitate screening of phase formability as a function of composition and temperature. However, the integration of computational predictions with experimental validation remains challenging in high-throughput studies. In this work, we introduce a quantitative confidence metric to assess the agreement between predictions and experimental observations, providing a quantitative measure of the confidence of machine learning models trained on either DFT or CALPHAD input in accounting for experimental evidence. The experimental dataset was generated via high-throughput in-situ synchrotron X-ray diffraction on compositionally varied FeNiMnCr alloy libraries, heated from room temperature to ~1000 °C. Agreement between the observed and predicted phases was evaluated using either temperature-independent phase classification or a model that incorporates a temperature-dependent probability of phase formation. This integrated approach demonstrates where strong overall agreement between computation and experiment exists, while also identifying key discrepancies, particularly in FCC/BCC predictions at Mn-rich regions to inform future model refinement.

36 - MATERIALS SCIENCE↗

Machine learning for domain transfer between simulated and experimental 2D X-ray diffraction patterns using generative adversarial networks

X-ray diffraction (XRD) is a well-established technique for analyzing materials at an atomic level. Dynamic compression experiments (DCE), in which materials are subject to extreme pressures, can provide fundamental understanding to pressure-induced phase transitions and compression of the crystal lattice. The analysis of XRD patterns from highly compressed samples is non-trivial given the sparsity of data, high experimental costs, and the fact that the data is often marred with X-ray background and other artifacts. While accurate computational frameworks exist, they solve the forward problem—from structures and orientations to XRD patterns. Solving the inverse problem for 2D experimental diffraction patterns is currently a complex manual process of matching and comparing experimentally observed patterns to computationally generated ones. Machine learning is a promising tool for automating the matching process but often requires data-intensive architectures. Here, in this study, we use a CycleGAN to translate the domain of limited experimental data to a domain in which there is readily available simulated data. This domain shift allows data-intensive machine learning models that have only been trained on simulated XRD patterns to be used in the analysis of experiments.

Brozak, Samantha Jean [Sandia National Laboratorie↗