Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning potentials”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Enriching the Twitter Stream Increasing Data Mining Yield and Quality Using Machine Learning

Social media data streams are important sources of real-time and historical global information for science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are exploring the Twitter data stream for its potential in augmenting the validation program of NASA Earth science missions, specifically the Global Precipitation Measurement (GPM) mission. We have implemented a tweet processing infrastructure that outputs classified precipitation tweets. Inputs are "passive" tweets, along with a smaller number of tweets from "active" participants, i.e., those knowingly contributing to our effort. The "active" tweets, presumably of higher quality, enrich the Twitter stream. "Active" sources include data scraped from other social media (e.g., public Facebook posts) and data from existing crowdsourcing programs (e.g., mPING reports). In addition, there is likely relevant precipitation information in images and documents that are the end points of links often included in tweets. Information derived from these "active" sources could then be tweeted into the Twitter stream, thus enriching its quality. The objective of our current work is to mine these tweet­ linked images and documents, using neural networks, to increase the information content and quality related to precipitation. For images, we classified them as either precipitation-related or not. For training and validation, we used images obtained via the Google custom search API. We created two models: (1) by training a simple Convolutional Neural Network and (2) by using transfer learning principles to adapt a pre-trained object recognition model. For documents, both those linked to tweets and the tweet contents, we trained Hierarchical Attention Networks to determine precipitation occurrence, type, and intensity. For training and validation, we used a keyword-filtered tweet data set labelled with ground truth data from Dark Sky (an API to retrieve weather-related labels) and the National Severe Storms Laboratory's Multi­ Radar/Multi-Sensor (MRMS) system. Our results demonstrated the efficacy of our machine learning approaches for enriching the Twitter stream, to derive information potentially useful for validation of earth science satellite data.

Albayrak, Arif↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗

Using Machine Learning to Develop a Predictive Model for Future Fire Seasons

The deep learning model shows promise for predicting areas of high wildfire potential. Full evaluation of the model performance is ongoing. Currently, the developed deep learning model is better overall at predicting the number of fires over the acres burned. Acres burned is dependent on location, suppression plan, and current conditions. Antecedent conditions are only one piece of the equation. In-season changes are not accounted for. An ignition source is required, which further complicates the model training and prediction.

White, Andrew T.↗

Seismic Elastic Double-Beam Characterization of Faults and Fractures for CO₂ Storage Site Selection

Site characterization for underground injection and storage of gigatonne-scale CO₂ requires reliable and cost-effective methods to detect and characterize faults and fractures and to assess their stress state and fault activation potential. This is critical, as wastewater injection and disposal have been shown to activate faults and induce earthquakes, and CO₂ leakage remains a key concern for long-term storage. In this project, we developed seismic methods to detect and characterize large-scale sedimentary and crystalline basement faults and associated small-scale fractures below conventional seismic imaging resolution using multicomponent (9C) surface seismic data. Machine learning was used to automatically interpret large-scale faults, providing key information for estimating the maximum magnitude of potential induced earthquakes. High-fidelity imaging was achieved by exploiting redundancy across multiple elastic wave modes, where independent images from different modes and frequencies cross-validate each other. We also used our nonlinear signal comparison (NLSC) method for ground roll removal, improving data quality in complex near-surface conditions. The methods were validated using field data acquired in central Montana. Results show that basement faults extend into the sedimentary section and that small-scale fractures are widespread above the basement. The inferred stress orientation is consistent with regional stress data, and the estimated maximum induced earthquake magnitude is small (Mw ~2.3). The developed workflow provides a practical approach for fault and fracture characterization and for assessing induced seismicity and leakage risk. It is directly applicable to CO₂ storage site selection and to other subsurface systems.

02 PETROLEUM↗

Exploring Geothermal Potential of Great Basin Sub-Regions

The INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project aims to discover new, economically viable hidden geothermal systems in the Great Basin region by building on previous work in play fairway analysis and machine learning. A key objective of this project is to develop an exploration workflow to reduce geothermal exploration risks for hidden geothermal systems. A single preliminary play fairway workflow was developed from the assessment of the regional INGENIOUS geological, geophysical, and geochemical datasets. This workflow provided new preliminary predictive geothermal fairway maps for the INGENIOUS study area, which encompasses most of Nevada, western Utah, southern Idaho, southeastern Oregon, and easternmost California. However, a recent study (incorporating machine learning techniques) of a portion of Nevada identified four geologic domains and determined that the relative importance of individual datasets or features as indicators of geothermal potential may differ across these domains. The INGENIOUS study area includes a much larger and more geologically diverse region; therefore, additional geologic domains or sub-regions are expected. To assess the sub-regions in the INGENIOUS study area, principal component analysis and k-means clustering were applied. Preliminary results indicate that the INGENIOUS regional data cluster into groups that relate to different geologic domains in the Great Basin region. These include domains such as the Walker Lane, extensional western Great Basin region, broad lower strain region in the eastern Great Basin of western Utah and eastern Nevada, Quaternary volcanic fields, and the area adjacent to the Snake River Plain. These clusters are assessed to determine the key geologic drivers of the identified clusters. Understanding this variability can provide key insights for the exploration and characterization of hidden geothermal systems in the Great Basin region and could indicate the need to develop multiple geothermal conceptual models and play fairway workflows for the INGENIOUS study area.

exploration↗

Exploring Geothermal Potential of Great Basin Sub-Regions: Preprint

The INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project aims to discover new, economically viable hidden geothermal systems in the Great Basin region by building on previous work in play fairway analysis and machine learning. A key objective of this project is to develop an exploration workflow to reduce geothermal exploration risks for hidden geothermal systems. A single preliminary play fairway workflow was developed from the assessment of the regional INGENIOUS geological, geophysical, and geochemical datasets. This workflow provided new preliminary predictive geothermal fairway maps for the INGENIOUS study area, which encompasses most of Nevada, western Utah, southern Idaho, southeastern Oregon, and easternmost California. However, a recent study (incorporating machine learning techniques) of a portion of Nevada identified four geologic domains and determined that the relative importance of individual datasets or features as indicators of geothermal potential may differ across these domains. The INGENIOUS study area includes a much larger and more geologically diverse region; therefore, additional geologic domains or sub-regions are expected. To assess the sub-regions in the INGENIOUS study area, principal component analysis and k-means clustering were applied. Preliminary results indicate that the INGENIOUS regional data cluster into groups that relate to different geologic domains in the Great Basin region. These include domains such as the Walker Lane, extensional western Great Basin region, broad lower strain region in the eastern Great Basin of western Utah and eastern Nevada, Quaternary volcanic fields, and the area adjacent to the Snake River Plain. These clusters are assessed to determine the key geologic drivers of the identified clusters. Understanding this variability can provide key insights for the exploration and characterization of hidden geothermal systems in the Great Basin region and could indicate the need to develop multiple geothermal conceptual models and play fairway workflows for the INGENIOUS study area.

GEOTHERMAL ENERGY↗

Fracture Network Quantification during CO2 Injection

This is the presentation prepared for the ARMA 2025 (59th US Rock Mechanics/Geomechanics Symposium) Conference held in Santa Fe, New Mexico, June 8-11, 2025. Accurate mapping and quantification of these networks are essential to ensure the integrity of CO2 storage reservoirs, understand and reduce potential leakage, and maintain long-term environmental safety. This study presents a novel machine learning-driven approach, integrated with geomechanical analysis, to quantify fracture networks and assess their spatial distribution during CO2 injection. This paper combines microseismic monitoring data with principles of hydraulic diffusivity and geomechanical analysis to characterize reservoir scale fracture network. The novelty of our approach lies in its capacity to assimilate time-dependent pressure data and microseismicity into a cohesive framework, which not only identifies microseismic triggering fronts but also tracks fracture distribution during active injection. Besides, leveraging image log data and analysis our approach also provides another angle of the insights to solidate the fracture networks understanding and geomechanical impacts. Key results from our study include the detection of over 100 distinct fracture clusters across the injection site, with fracture orientations strongly correlated with the prevailing in-situ stress field.

CO2 storage and sequestration↗

MLtool++ package for machine learning and its applications to materials data

We are developing Mltool++ package of software programs for machine learning (ML). Given the MLtool Python code, we create a faster C++ code with the potential for parallelization. We have extracted materials data from the literature. One dataset contains melting temperatures of stoichiometric 1:1 metallic compounds XZ, composed by elements X={Al, Ti, V, Cr, Zr, Nb, Mo, Hf, Ta, W} and Z={Co, Ni, Cu, Rh, Pd, Ag, Ir, Pt, Au}, and another contains solid-solid symmetry-breaking phase transition temperatures. We studied dependences of temperatures on composition, found several correlations, and parametrized them by analytical functions. Mltool++ package is generic and applicable to any tabulated numeric data.

Pierce M. Pettit↗

LLMs and GenAI Tools to Depict Contributions of Human Systems to Spaceflight Tasks Execution

Recent advancements in Artificial Intelligence and Machine Learning (AI/ML) technologies, particularly Large Language Models (LLMs) capable of sophisticated syntax analysis, offer substantial potential in automating complex processes, thereby saving time and human resources. This study explores the development of an LLM-driven model designed to analyze and categorize a diverse set of Mars mission tasks into 18 predefined Human System Task Categories (HSTCs) based on their textual descriptions. As part of developing the Crew Health and Performance – Probabilistic Risk Assessment (CHP-PRA projects Performance Risk Model (PRisM) proof-of-concept, we established a framework to project performance scores from small-scale tests onto a preliminary list of Mars tasks. The foundation of our model was a comprehensive spreadsheet populated by NASA experts and clinicians, which detailed each Mars task alongside binary indicators of HSTC involvement. This dataset enabled the initial application of supervised ML, training and testing on existing HSTC labels. The HSTCs were originally defined from a medical system perspective, focusing on task impairments due to deteriorated human health. To expand our model's scope to include categories impacting performance, we face the challenge of generating binary labels (0 or 1) for new categories without pre-existing data. We address this by employing Generative AI (GenAI) software to determine whether a given task involved a new category by asking, "Does task A involve using category B?" We validate our approach by comparing the GenAI's binary classifications with the expert-provided labels for existing HSTCs. Notably, we utilize Ollama [4], a locally hosted GenAI tool that does not require cloud access, thus safeguarding NASA's proprietary data from unauthorized exposure. This study demonstrates the feasibility of leveraging cutting-edge AI tools to advance research, paving the way for automation and rapid decision-making in space exploration.

Mona Matar↗

Fracture Network Quantification during CO2 Injection

This is the conference paper accompanying an oral presentation at the ARMA 2025 (59th US Rock Mechanics/Geomechanics Symposium) Conference held in Santa Fe, New Mexico, June 8-11, 2025. Accurate mapping and quantification of these networks are essential to ensure the integrity of CO2 storage reservoirs, understand and reduce potential leakage, and maintain long-term environmental safety. This study presents a novel machine learning-driven approach, integrated with geomechanical analysis, to quantify fracture networks and assess their spatial distribution during CO2 injection. This paper combines microseismic monitoring data with principles of hydraulic diffusivity and geomechanical analysis to characterize reservoir scale fracture network. The novelty of our approach lies in its capacity to assimilate time-dependent pressure data and microseismicity into a cohesive framework, which not only identifies microseismic triggering fronts but also tracks fracture distribution during active injection. Besides, leveraging image log data and analysis our approach also provides another angle of the insights to solidate the fracture networks understanding and geomechanical impacts. Key results from our study include the detection of over 100 distinct fracture clusters across the injection site, with fracture orientations strongly correlated with the prevailing in-situ stress field.

CO2 storage and sequestration↗

Development of a machine learning model for polyethylene pyrolysis using a detailed reaction mechanism

Waste plastics have recently received significant attention as the issue of waste generation continues to increase. Thermal conversion processes, such as pyrolysis and gasification, are attractive potential technologies for utilizing waste plastics and reducing overall waste generation. Efficient utilization of plastics requires a detailed understanding of the conversion process such as pyrolysis and gasification. However, a mechanistic understanding of these processes lead to large and complex kinetic schemes that are not suited for large-scale and long-time simulation methods. Currently, most modeling approaches for pyrolysis and gasification rely on globally lumped, simplified kinetic schemes that provide results that are classified by their product type and not individual species, which limit the level of fidelity achieved via modeling. A machine learning (ML) model has been developed for the primary reactions of high-density polyethylene (HDPE) in an attempt to increase computational efficiency while still maintaining a high level of detail and accuracy. The ML model is trained on a detailed reaction mechanism containing 42 total species and 737 chemical reactions. A DeepONet branch and trunk architecture was adopted to train the model using time-steps relevant to computational fluid dynamics simulations. The ML used physics-informed loss functions to ensure mass conservation. The surrogate model has been deployed in simple MFiX CFD simulations, single particle and an experimental drop tube reactor, and has shown promising performance compared to the original scheme.

Houston, Ross↗

Tuning water dissociation at oxide–electrolyte interfaces with electric fields

Understanding how electric fields influence water dissociation at heterogeneous interfaces is crucial for controlling interfacial chemical reactions and advancing next-generation energy technologies. Herein, ab initio–based machine learning simulations show that even small electric field changes can significantly alter the water dissociation fraction at planar TiO 2 –electrolyte interfaces. The resulting free energy difference between undissociated and dissociated interfacial water exhibits a linear dependence on the field change with a slope of 1.97 eÅ, which far exceeds the dissociation-induced dipole change of a water molecule. Employing a machine-learned collective variable to investigate the reaction statistics of thousands of water dissociation/recombination events, we find that small electric field changes exert minor effects on individual reaction energy barriers but significantly influence the populations of local configurations associated with initial states that are most favorable for reactions. These findings elucidate the pronounced impact of electric fields on interfacial water dissociation and reveal a mechanism for electric-field-controlled chemical reactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Revisiting the Ground Magnetic Field Perturbations Challenge: A Machine Learning Perspective

Forecasting ground magnetic field perturbations has been a long-standing goal of the space weather community. The availability of ground magnetic field data and its potential to be used in geomagnetically induced current studies, such as risk assessment, have resulted in several forecasting efforts over the past few decades. One particular community effort was the Geospace Environment Modeling (GEM) challenge of ground magnetic field perturbations that evaluated the predictive capacity of several empirical and first principles models at both mid- and high-latitudes in order to choose an operative model. In this work, we use three different deep learning models-a feed-forward neural network, a long short-term memory recurrent network and a convolutional neural network-to forecast the horizontal component of the ground magnetic field rate of change (dB H /dt) over 6 different ground magnetometer stations and to compare as directly as possible with the original GEM challenge. We find that, in general, the models are able to perform at similar levels to those obtained in the original challenge, although the performance depends heavily on the particular storm being evaluated. We then discuss the limitations of such a comparison on the basis that the original challenge was not designed with machine learning algorithms in mind.

Victor A. Pinto↗

Detecting Process Equipment Failures Using Acoustic Data and Machine Learning

Nuclear power plant (NPP) process equipment such as fans, motors, valves, and pumps generate frequent or continuous noise, and deviations from the normal operational sounds made by this equipment can indicate potential issues. These deviations can be identified via automated acoustic anomaly detection, which involves using acoustic sensors (i.e., microphones) alongside detection algorithms to continuously monitor for changes in acoustic signatures. This task is made challenging by the substantial background noise that exists, such as operators opening and closing doors, manipulating valves, and conversing—in addition to typical plant noises. In collaboration with a nuclear power utility partner, this effort assessed the efficacy of acoustic anomaly detection when using a specific acoustic sensor that compresses data into a fixed set of features that are transferable over a standard Internet of Things communication protocol, thereby improving usability but potentially degrading detection performance. Two methods of performing automated acoustic anomaly detection were evaluated: one-class support vector machine (OC-SVM) and isolation forest (iForest). To enable the use of high-quality acoustic data encompassing both normal and anomalous conditions, the study utilized the publicly available Malfunctioning Industrial Machine Investigation and Inspection dataset, which includes real measured acoustic sensor data for a range of equipment types, model numbers, and signal-to-noise ratios (SNRs), along with a benchmark set of detection results. Using this dataset, the methods were tested and then compared against the benchmark results. The results indicated that although the specific acoustic sensor did not enable as rich a feature set extraction, the proposed methods with the limited feature set performed just as well. This provides solid justification for both the methods and the use of the proposed acoustic sensor.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Utilizing Machine Learning to Improve Neutralization Potency of an HIV-1 Antibody Targeting the gp41 N-Heptad Repeat

The N-heptad repeat (NHR) of the HIV-1 gp41 prehairpin intermediate (PHI) is an attractive potential vaccine target with high sequence conservation across diverse strains. However, despite the potency of NHR-targeting peptides and clinical efficacy of the NHR-targeting entry inhibitor enfuvirtide, no potently neutralizing NHR-directed monoclonal antibodies (mAbs) nor antisera have been identified or elicited to date. The lack of potent NHR-binding mAbs both dampens enthusiasm for vaccine development efforts at this target and presents a barrier to performing passive immunization experiments with NHR-targeting antibodies. To address this challenge, we previously developed an improved variant of the NHR-directed mAb D5, called D5_AR, which is capable of neutralizing diverse tier-2 viruses. Building on that work, here we present the 2.7Å-crystal structure of D5_AR bound to NHR mimetic peptide IQN17. We then utilize protein language models and supervised machine learning to generate small (n < 100) libraries of D5_AR variants that are subsequently screened for improved neutralization potency. We identify a variant with 5-fold improved neutralization potency, D5_FI, which is the most potent NHR-directed monoclonal antibody characterized to date and exhibits broad neutralization of tier-2 and −3 pseudoviruses as well as replicating R5 and X4 challenge strains. Additionally, our work highlights the ability of protein language models to efficiently identify improved mAb variants from relatively small libraries.

Biopolymers↗

Anomaly Detection in Electronic Health Records Across Hospital Networks: Integrating Machine Learning With Graph Algorithms

In a large hospital system, a network of hospitals relies on electronic health records (EHRs) to make informed decisions regarding their patients in various clinical domains. Consequently, the dependability of the health information technology (HIT) systems responsible for collecting EHR data is of utmost importance for patient safety. Recently, novel methods and tools aimed at identifying anomalies in EHR data to bolster the reliability of HIT systems have been introduced. However, these existing methods and tools primarily concentrate on individual hospitals, which limits our understanding of system-wide anomalous events and their potential impact on patient safety across multiple hospitals. In this article, we introduce a new approach to detecting anomalies in EHR data within a network of hospitals. This is achieved by combining advanced machine learning techniques with graph algorithms to create a tool capable of swiftly identifying and responding to deviations. Our proposed approach employs a combination of five machine learning models, harnessing the unique strengths of each model to provide a more robust detection system. The detected anomalies are then represented as graphs, allowing us to recognize patterns across the hospital network. This aids in identifying anomalies that span multiple medical facilities, potentially indicating broader system-level risks. Extensive real-world testing of our approach demonstrated its ability to offer actionable insights compared to existing methods. Additionally, its scalable design ensures seamless integration into existing HIT infrastructures.

Niu, Haoran [Oak Ridge National Laboratory (ORNL),↗

Evaluation of GlassNet for physics-informed machine learning of glass stability and glass-forming ability

Glassy materials form the basis of many modern applications, including nuclear waste immobilization, touch-screen displays, and optical fibers, and also hold great potential for future medical and environmental applications. However, their structural complexity and large composition space make design and optimization challenging for certain applications. Of particular importance for glass processing and design is an estimate of a given composition's glass-forming ability (GFA). However, there remain many open questions regarding the underlying physical mechanisms of glass formation, especially in oxide glasses. It is apparent that a proxy for GFA would be highly useful in glass processing and design, but identifying such a surrogate property has proven itself to be difficult. While glass stability (GS) parameters have historically been used as a GFA surrogate, recent research has demonstrated that most of these parameters are not accurate predictors of the GFA of oxide glasses. Here, in this work, we explore the application of an open-source pre-trained neural network model, GlassNet, that can predict the characteristic temperatures necessary to compute GS with reasonable performance and assess the feasibility of using these physics-informed machine learning (PIML)-predicted GS parameters to estimate GFA. In doing so, we track the uncertainties at each step of the computation—from the original ML prediction errors to the compounding of errors during GS estimation, and finally to the final estimation of GFA. While GlassNet exhibits reasonable accuracy on all individual properties, we observe a large compounding of error in the combination of these individual predictions for the PIML prediction of GS, finding that random forest models offer similar accuracy to GlassNet. We also break down the performance of GlassNet on different glass families and find that the error in GS prediction is correlated with the error in crystallization peak temperature prediction. Lastly, we utilize this finding to assess the relationship between top-performing GS parameters and GFA for two ternary glass systems: sodium borosilicate and sodium iron phosphate glasses. We conclude that to obtain true ML predictive capability of GFA, significantly more data needs to be collected.

36 MATERIALS SCIENCE↗

Avoiding Braess' Paradox Through Collective Intelligence

In an Ideal Shortest Path Algorithm (ISPA), at each moment each router in a network sends all of its traffic down the path that will incur the lowest cost to that traffic. In the limit of an infinitesimally small amount of traffic for a particular router, its routing that traffic via an ISPA is optimal, as far as cost incurred by that traffic is concerned. We demonstrate though that in many cases, due to the side-effects of one router's actions on another routers performance, having routers use ISPA's is suboptimal as far as global aggregate cost is concerned, even when only used to route infinitesimally small amounts of traffic. As a particular example of this we present an instance of Braess' paradox for ISPA'S, in which adding new links to a network decreases overall throughput. We also demonstrate that load-balancing, in which the routing decisions are made to optimize the global cost incurred by all traffic currently being routed, is suboptimal as far as global cost averaged across time is concerned. This is also due to "side-effects", in this case of current routing decision on future traffic. The theory of COllective INtelligence (COIN) is concerned precisely with the issue of avoiding such deleterious side-effects. We present key concepts from that theory and use them to derive an idealized algorithm whose performance is better than that of the ISPA, even in the infinitesimal limit. We present experiments verifying this, and also showing that a machine-learning-based version of this COIN algorithm in which costs are only imprecisely estimated (a version potentially applicable in the real world) also outperforms the ISPA, despite having access to less information than does the ISPA. In particular, this COIN algorithm avoids Braess' paradox.

Wolpert , David H.↗