Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Characterization and prediction of the electromechanical wear of contact tips during wire arc additive manufacturing of 316L stainless steel

Here, this study seeks to better understand the degradation of the contact tip with respect to WAAM for a 316L wire electrode as well as explore methods of monitoring the contact tip state from process data. The contact tip, a consumable component, positions the wire and serves as the electrical contact surface between the wire electrode and the welding power supply. The wear of the contact tip was characterized in terms of material loss and material contamination for a set of tips worn to discrete levels as measured by the amount of wire fed or arc time. Geometrical characterization found a 49% increase in the bore exit area at 180 meters of wire fed. Machine learning models were developed to predict the relative bore exit area of the contact tip from arc-based process data and a random forest classifier exhibited favorable performance with a cross-validated f1-score of 0.84. The regression architecture implemented a multi-layer perceptron with the ability to predict the relative exit area with an $R^2$ score of 0.75. Key features used in the prediction include the standard deviation of the voltage and the time between shorts.

Contact tip wear↗

Projected increases in tropical cyclone-induced U.S. electric power outage risk

Abstract While power outages caused by tropical cyclones (TCs) already pose a great threat to coastal communities, how—and why—these risks will change in a warming climate is poorly understood. To address this need, we develop a robust machine learning model to capture TC-induced power outage risk. When applied to 900 000 synthetic TCs downscaled from simulated historical and future climate conditions under a strong warming scenario, we find outage risk in the United States and Puerto Rico is expected to increase broadly by the end of the century, with some states seeing increases of 60% and higher. Further, we discover that rising rainfall rates will play an increasingly important role in TC-induced power outage risk as the climate changes, explaining more than 50% of the projected change in risk in some regions. These insights are important for guiding decision-makers in their future outage risk investment and mitigation plans.

Grid Resilience↗

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields in the last two decades, in part thanks to an increasing culture of open data sharing and reuse. Due to its capability for identifying complex relationships and patterns, AI/ML methodology is particularly well suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are many key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Even with the positive culture of Open Science and data sharing, inexperienced researchers working quickly without proper checks can produce models that perform poorly outside of the immediate training dataset. Lessons learned from biological AI/ML research indicate that Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Andrew Casaletto↗

Learning and discovering multiple solutions using physics-informed neural networks with random initialization and deep ensemble

In this work we explore the capability of physics-informed neural networks (PINNs) to discover multiple solutions. Many real-world phenomena governed by nonlinear differential equations (DEs), such as fluid flow, exhibit multiple solutions under the same conditions, yet capturing this solution multiplicity remains a significant challenge. A key difficulty lies in providing appropriate initial conditions or guesses, as widely used time-marching schemes and Newton’s method are highly sensitive to these choices when solving complex computational problems. While machine learning models, particularly PINNs, have shown promise in solving DEs, their ability to capture multiple solutions remains underexplored. In this work, we propose a simple and practical approach using PINNs to learn and discover multiple solutions. We first demonstrate that PINNs, when combined with random initialization and deep ensemble method—originally developed for uncertainty quantification—can effectively uncover multiple solutions to nonlinear ordinary and partial DEs. Although training large ensembles of PINNs may appear computationally demanding, this can be done efficiently using vectorization techniques supported by modern deep learning frameworks, allowing many networks to be trained simultaneously. Our approach highlights the critical role of initialization in shaping solution diversity, addressing an often-overlooked aspect of machine learning for scientific computing. Furthermore, we propose utilizing PINN-generated solutions as initial conditions or initial guesses for conventional numerical solvers to enhance accuracy and efficiency in capturing multiple solutions. Extensive numerical experiments, including the Allen–Cahn equation and cavity flow, where our approach successfully identifies both stable and unstable solutions, validate the effectiveness of our method. These findings establish a general and efficient framework for addressing solution multiplicity in nonlinear DEs.

97 MATHEMATICS AND COMPUTING↗

Secondary organic aerosols derived from intermediate-volatility n-alkanes adopt low-viscous phase state

Abstract. Secondary organic aerosol (SOA) derived from n-alkanes, as emitted from vehicles and volatile chemical products, is a major component of anthropogenic particulate matter, yet the chemical composition and phase state are poorly understood and thus poorly constrained in aerosol models. Here we provide a comprehensive analysis of n-alkane SOA by explicit gas-phase chemistry modeling, machine learning, and laboratory experiments to show that n-alkane SOA adopts low-viscous semi-solid or liquid states. Our study underlines the complex interplay of molecular composition and SOA viscosity: n-alkane SOA with a higher carbon number mostly consists of less functionalized first-generation products with lower viscosity, while the SOA with a lower carbon number contains more functionalized multigenerational products with higher viscosity. This study opens up a new avenue for analysis of SOA processes, and the results indicate few kinetic limitations of mass accommodation in SOA formation, supporting the application of equilibrium partitioning for simulating n-alkane SOA formation in large-scale atmospheric models.

54 ENVIRONMENTAL SCIENCES↗

High temperature melting of dense molecular hydrogen from machine-learning interatomic potentials trained on quantum Monte Carlo

We present results and discuss methods for computing the melting temperature of dense molecular hydrogen using a machine learned model trained on quantum Monte Carlo data. In this newly trained model, we emphasize the importance of accurate total energies in the training. We integrate a two phase method for estimating the melting temperature with estimates from the Clausius–Clapeyron relation to provide a more accurate melting curve from the model. We make detailed predictions of the melting temperature, solid and liquid volumes, latent heat, and internal energy from 50 to 180 GPa for both classical hydrogen and quantum hydrogen. At pressures of roughly 173 GPa and 1635 K, we observe molecular dissociation in the liquid phase. Here, we compare with previous simulations and experimental measurements.

08 HYDROGEN↗

Do Better Satellite Precipitation Algorithms Improve Landslide Hazard Assessment?

Satellites make it possible to estimate precipitation in near real time. Given the challenges of achieving global coverage by other means, these data are used widely. However, few systems for landslide hazard assessment rely on satellite precipitation estimates. This could be due in part to perceptions of accuracy, although latency, spatial resolution, and other factors may also be important. We test whether recent changes to data streams from the Global Precipitation Measurement mission (GPM) have improved its potential for use in landslide prediction. Specifically, we examine data produced by the Integrated Multi-satellitERetrievals for the GPM (IMERG) algorithm, which was upgraded to version 7 this year. IMERG relies upon other algorithms, including the Goddard Profiling Algorithm (GPROF) and the GPM Combined Radar-Radiometer Algorithm (CORRA). Many changes have been made during the switch from IMERG version 6 to version 7. These include upgrading CORRA and GPROF to version 7, to improve the accuracy of precipitation in frozen, mountainous, and coastal areas. The measured intensity of some storms has been enhanced with a new algorithm, the Scheme for Histogram Adjustment with Ranked Precipitation Estimates in the Neighborhood. Combined with many others, these changes to IMERG should improve its utility for landslide hazard assessment in a variety of contexts. To test this idea, we retrain the global Landslide Hazard Assessment for Situational Awareness (LHASA) model twice—first with data from IMERG version 6B and second with 7B. Since current daily rainfall is the most important variable in determining outcomes predicted by LHASA, it should reflect changes made to that input. First, we grid the landslides at a daily, thirty-arcsecond resolution. This serves as the response variable. At each of these sites current and antecedent rainfall are extracted, along with antecedent snow mass and soil moisture, slope, and PGA. In addition, one million grid cells are selected at random points to represent conditions under which landslides (probably) do not occur. After merging these data, we hold back 20% of the dataset for validation purposes and train a machine-learning model with the rest. We assess both the model’s overall ability to identify landslides and its ability to predict specific large landslide disasters.

Thomas A Stanley↗

A Gaussian process based surrogate approach for the optimization of cylindrical targets

Simulating direct-drive inertial confinement experiments presents significant computational challenges, both due to the complexity of the codes required for such simulations and the substantial computational expense associated with target design studies. Machine learning models, and in particular, surrogate models, offer a solution by replacing simulation results with a simplified approximation. In this study, we apply surrogate modeling and optimization techniques that are well established in the existing literature to one-dimensional simulation data of a new cylindrical target design containing deuterium–tritium fuel. These models predict yields without the need for expensive simulations. We find that Bayesian optimization with Gaussian process surrogates enhances sampling efficiency in low-dimensional design spaces but becomes less efficient as dimensionality increases. Nonetheless, optimization routines within two-dimensional and five-dimensional design spaces can identify designs that maximize yield, while also aligning with established physical intuition. Optimization routines, which ignore constraints on hydrodynamic instability growth, are shown to lead to unstable designs in 2D, resulting in yield loss. However, routines that utilize 1D simulations and impose constraints on the in-flight aspect ratio converge on novel cylindrical target designs that are stable against hydrodynamic instability growth in 2D and achieve high yield.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Advancing stream temperature prediction with a generalizable large-sample framework across CONUS river reaches

Accurately predicting stream temperature in ungauged basins remains a critical challenge for water resource management, thermoelectric power plant cooling, and ecosystem conservation. Large-sample machine learning models trained on hundreds of well-monitored river basins have shown remarkable performance; however, such models have yet to be developed solely using forcing data that can be readily extracted to simulate stream temperatures anywhere in the contiguous United States (CONUS). In this study, we present a scalable, large-sample deep learning framework using Long Short-Term Memory (LSTM) networks to simulate daily stream temperatures in ungauged basins across the CONUS. The framework leverages both modeled reanalysis of meteorological and streamflow inputs as well as static attributes available for all 2.7 million CONUS river reaches in the National Hydrography Dataset Plus (NHDPlusV2). By generating dynamical inputs from predefined thermally relevant upstream contributing areas, rather than the entire upstream basin, the model also offers improvements in very large basins where full-basin averaging can dilute the most important influences on stream temperature. Evaluated across 300 basins, the model achieves a median Mean Absolute Error (MAE) of 1.1 °C and a Nash-Sutcliffe Efficiency (NSE) of 0.95 on temporally and spatially distinct test folds—comparable to models trained exclusively using meteorological and streamflow observational data. The flexible, high-performing framework generalizes to any unmonitored river reach without significant regulation or unnatural thermal input immediately upstream, substantially expanding predictive capabilities in data-scarce regions.

Hydrology↗

Sensor Reduction for Diversion Detection in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors are designed as a smaller, cheaper, and safer alternative to traditional nuclear power plants. Their non-traditional characteristics and prospect of mass production and deployment will likely require new approaches to nuclear safeguards. The primary proliferation concern with microreactors is the diversion of fuel material. Such diversion may produce measurable defects in key physical attributes like neutron flux, which may in turn be detectable using machine learning models. Preliminary work has demonstrated this ability for modeled nominal and diversion scenarios using large quantities of energy integrated neutron flux data. In practice, the number of available sensors for such measurements will be limited and energy integrated flux information will not be available. This work explores the ability of tree-based gradient boosted ensemble models to classify a given microreactor core is nominal or diversion, and determine the number of fuel pins diverted in the case of diversion with reduced numbers of sensors and more realistic detector responses. Classification accuracy of greater than 98% and regression errors as low as 5% of the total number of fuel pins were achieved with as few as 15 sensors, compared to 99% and 4.1% with a maximum of 240 sensors.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Investigating the opioid epidemic across the United States: Associations between county-level characteristics and overdose mortality

The opioid crisis remains a critical public health challenge in the United States. Despite national efforts that reduced opioid prescribing by nearly 44% between 2011 and 2021, opioid overdose deaths more than tripled during the same period. This alarming trend reflects a major shift in the crisis, with illegal opioids now driving the majority of overdose deaths instead of prescription opioids. Although supply-side factors fueling this transition have been widely studied, the structural and community-level conditions that shape overdose mortality are less well understood. To help address this gap, this study has three primary objectives: (1) overcome structural gaps in national data to construct a complete nationwide county-level dataset from 2010 to 2022; (2) using data analysis, identify and investigate spatiotemporal anomalies in overdose mortality; and (3) using two machine-learning models, quantify the importance of thirteen social vulnerability variables in predicting overdose mortality. Our results identify unemployment and limited vehicle access as key county-level predictors of overdose mortality. Higher levels of these vulnerabilities are associated with elevated mortality, whereas lower levels are associated with reduced mortality. These findings highlight factors that may be relevant for public health planning and policy prioritization within the context of the opioid crisis.

Anomaly analysis↗

Fundamental limit of jet tagging

Identifying the origin of high-energy hadronic jets (jet tagging) has been a critical benchmark problem for machine learning in particle physics. Jets are ubiquitous at colliders and are complex objects that serve as prototypical examples of collections of particles to be categorized. Over the last decade, machine learning-based classifiers have replaced classical observables as the state of the art in jet tagging. Increasingly complex machine learning models are leading to increasingly more effective tagger performance. Our goal is to address the question of convergence—are we getting close to the fundamental limit on jet tagging or is there still potential for computational, statistical, and physical insights for further improvements? We address this question using state-of-the-art generative models to create a realistic, synthetic dataset with a known jet tagging optimum. Various state-of-the-art taggers are deployed on this dataset, showing that there is a significant gap between their performance and the optimum. Our dataset and software are made public to provide a benchmark task for future developments in jet tagging and other areas of particle physics.

Artificial intelligence↗

Classification of Cloud Particle Imagery and Thermodynamics (COCPIT): A New Databasing Tool for the Characterization of Cloud Particle Images Captured During DOE Field Campaigns

The Department of Energy for decades has explored the earth system and atmosphere through research and deployment of in-situ and remote sensing platforms during field campaigns. Among these datasets exists a vast supply of cloud particle images that provide visual insight into the complex microphysics in the clouds that span our globe. The millions of images collected over decades of deployments provides a unique opportunity to further our understanding of our atmosphere down to the crystal size. This work over the past 5 years has sought to organize these images into digestible datasets that can then be used by scientists to further our understanding of microphysics. A machine learning model was developed that categorizes over 1.5 million images across 11 weather events with over 90% accuracy according to particle type. The database was then extended to include dimensional characteristics of the particle as well as co-location of environmental properties, such as temperature and water content. Then, to initialize the connection between these data and our understanding of how crystals form and grow, weather research and forecasting simulations were run to generate the growth histories of the classified crystals. This research culminates with 2 databases per event: (1) a database of all classified crystals and their dimensional and environmental properties and (2) simulated growth histories of each crystal. Finally, a user interface was created to allow researchers to explore data statistics.

54 ENVIRONMENTAL SCIENCES↗

Imbalanced Multi-layer Cloud Classification with Advanced Baseline Imager (ABI) and CloudSat/CALIPSO Data

Clouds at different altitudes play different roles in Earth’s climate. Comprehensive understanding of overlapping clouds is important for climate and weather prediction. The East Pacific region is where El Ni˜no and La Ni˜na originate and where multi-layer clouds frequently occur. The overlap of clouds at different altitudes in this region increases the classification complexity for cloud-based climatological studies. Unlike prior work in cloud layer classification that assumes single layer or two-layer of clouds, in this work, we consider multi-layer cloud classification with 8 cloud-level classes (clear-sky, high, middle, low, high+middle, high+low, middle+low, high+middle+low). We develop and analyze machine learning models on features extracted from satellite images from the East Pacific regions collected by GOES Advanced Baseline Imager (ABI). These are used to classify CloudSat/CALIPSO observed multi-layer clouds. Due to the imbalanced nature of the data, we investigate the adoption of conventional resampling methods, as well as deep learning methods with data augmentation. In our experiments, we utilize the random forest classifier and Multilayer perceptron classifier with data augmentation methods to reduce the class imbalance during training. With these approaches, we achieve a classification accuracy of 83.6% without exploiting any ancillary information.

machine learning↗

Satellite Based Precipitation Estimation in Orographic Regions within the Southwestern United States

Predicting precipitation-induced landslides requires accurate estimation of orographic precipitation. Many research studies have been done to estimate orographic precipitation using measurement/estimation methods that include rain gauges, ground-based radar, satellite-based estimates, and modeling. Each method has strengths and weaknesses, but none have been able to fully resolve orographic precipitation. The Integrated Multi-satellitE Retrievals for Global Precipitation Measurement Mission (IMERG) early run product provides precipitation estimates at a 0.1° spatial resolution, at half-hour time scales, with only a 4-hour latency. This product is currently used for the global landslide hazard assessment, but there are known issues in mountainous terrain that the IMERG algorithm has not been able to fully resolve; and the sparse gauge density in these regions makes it even more difficult. In this study, precipitation events were identified using high temporal resolution (5-15 minute) precipitation observations from rain gauges in mountainous terrain in the southwestern United States. The brightness temperature from several infrared (IR) bands from the Geostationary Operational Environmental Satellite (GOES) 16 satellite were used to estimate precipitation with a k-nearest neighbor machine learning model. Compared to IMERG, the IR estimates from GOES-16 performed better at predicting gauge-identified precipitation events. Additionally, the IR-only based estimates were able to estimate precipitation when IMERG failed to detect precipitation, false negative events. While additional analysis is needed, results indicate the need for better integration of IR observations for more accurate precipitation estimation in mountainous regions.

Jessica Sutton↗

Low-Cost Sensor Performance Intercomparison, Correction Factor Development, and 2+ Years of Ambient PM2.5 Monitoring in Accra, Ghana

Particulate matter air pollution is a leading cause of global mortality, particularly in Asia and Africa. Addressing the high and wide-ranging air pollution levels requires ambient monitoring, but many low- and middle-income countries (LMICs) remain scarcely monitored. To address these data gaps, recent studies have utilized low-cost sensors. These sensors have varied performance, and little literature exists about sensor intercomparison in Africa. By colocating 2 QuantAQ Modulair-PM, 2 PurpleAir PA-II SD, and 16 Clarity Node-S Generation II monitors with a reference-grade Teledyne monitor in Accra, Ghana, we present the first intercomparisons of different brands of low-cost sensors in Africa, demonstrating that each type of low-cost sensor PM2.5 is strongly correlated with reference PM2.5, but biased high for ambient mixture of sources found in Accra. When compared to a reference monitor, the QuantAQ Modulair-PM has the lowest mean absolute error at 3.04 μg/m3, followed by PurpleAir PA-II (4.54 μg/m3) and Clarity Node-S (13.68 μg/m3). We also compare the usage of 4 statistical or machine learning models (Multiple Linear Regression, Random Forest, Gaussian Mixture Regression, and XGBoost) to correct low-cost sensors data, and find that XGBoost performs the best in testing (R2: 0.97, 0.94, 0.96; mean absolute error: 0.56, 0.80, and 0.68 μg/m3 for PurpleAir PA-II, Clarity Node-S, and Modulair-PM, respectively), but tree-based models do not perform well when correcting data outside the range of the colocation training. Therefore, we used Gaussian Mixture Regression to correct data from the network of 17 Clarity Node-S monitors deployed around Accra, Ghana, from 2018 to 2021. We find that the network daily average PM2.5 concentration in Accra is 23.4 μg/m3, which is 1.6 times the World Health Organization Daily PM2.5 guideline of 15 μg/m3. While this level is lower than those seen in some larger African cities (such as Kinshasa, Democratic Republic of the Congo), mitigation strategies should be developed soon to prevent further impairment to air quality as Accra, and Ghana as a whole, rapidly grow.

Humidity↗

Establishing nationwide power system vulnerability index across US counties using interpretable machine learning

Power outages have become increasingly frequent, intense, and prolonged in the US due to climate change, aging electrical grids, and rising energy demand. However, largely due to the absence of granular spatiotemporal outage data, we lack data-driven evidence and analytics-based metrics to quantify power system vulnerability. This limitation has hindered the ability to effectively evaluate and address vulnerability to power outages in US communities. Here, in this work, we collected ∼179 million power outage records at 15-min intervals across 3022 US contiguous counties (96.15 % of the area) from 2014 to 2023. We developed a power system vulnerability assessment framework based on three dimensions (intensity, frequency, and duration) and applied interpretable machine learning models (XGBoost and SHAP) to compute Power System Vulnerability Index (PSVI) at the county level. Our analysis reveals a consistent increase in power system vulnerability across the US counties over the past decade. We identified 318 counties across 45 states as hotspots for high power system vulnerability, particularly in the West Coast (California and Washington), the East Coast (Florida and the Northeast area), the Great Lakes megalopolis (Chicago-Detroit metropolitan areas), and the Gulf of Mexico (Texas). Our heterogeneity analysis indicates that urban counties and those located along regional transmission boundaries tend to exhibit significantly higher vulnerability. Our results highlight the significance of the proposed PSVI for evaluating the vulnerability of communities to power outages. The findings underscore the widespread and pervasive impact of power outages across the country and offer crucial insights to support infrastructure operators, policymakers, and emergency managers in formulating policies and programs aimed at enhancing the resilience of the US power infrastructure.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Critical statistical assessment of data in metal additive manufacturing

Obtaining high quality data reflecting the relationships between the additive manufacturing (AM) process parameters, material microstructure and mechanical properties is crucial for the use of machine learning in AM. A database of over 4,000 data entries of metal AM was created thanks to a large number of literature studies on key process parameters and indicators of build quality. Meta-analysis reveals critical biases in the literature. Firstly, majority of studies report only high quality builds, these imbalances in reporting result in weak correlation between process parameters, properties and consolidation, limiting the ability of machine learning models to generalize beyond optimized conditions. Nevertheless, the trained models accurately predict yield strength ($R^2 = 0.85$), suggesting that certain process–property relationships are effectively captured within these models. Secondly, quantitative microstructural data are largely absent, limiting the learning of the microstructure-mechanical properties relationships. Finally, current process window identification is based largely on the consolidation, despite significant uncertainty in its measurement. It is important to identify the process map on the basis of not only the consolidation, but also mechanical behaviour under loading. Such a identification shows that 316 L and Inconel have much larger process map (i.e. highly printable) in comparison to the AlSi10Mg and Ti6Al4V.

Additive manufacturing↗