Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning potential”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Tuning water dissociation at oxide–electrolyte interfaces with electric fields

Understanding how electric fields influence water dissociation at heterogeneous interfaces is crucial for controlling interfacial chemical reactions and advancing next-generation energy technologies. Herein, ab initio–based machine learning simulations show that even small electric field changes can significantly alter the water dissociation fraction at planar TiO 2 –electrolyte interfaces. The resulting free energy difference between undissociated and dissociated interfacial water exhibits a linear dependence on the field change with a slope of 1.97 eÅ, which far exceeds the dissociation-induced dipole change of a water molecule. Employing a machine-learned collective variable to investigate the reaction statistics of thousands of water dissociation/recombination events, we find that small electric field changes exert minor effects on individual reaction energy barriers but significantly influence the populations of local configurations associated with initial states that are most favorable for reactions. These findings elucidate the pronounced impact of electric fields on interfacial water dissociation and reveal a mechanism for electric-field-controlled chemical reactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Revisiting the Ground Magnetic Field Perturbations Challenge: A Machine Learning Perspective

Forecasting ground magnetic field perturbations has been a long-standing goal of the space weather community. The availability of ground magnetic field data and its potential to be used in geomagnetically induced current studies, such as risk assessment, have resulted in several forecasting efforts over the past few decades. One particular community effort was the Geospace Environment Modeling (GEM) challenge of ground magnetic field perturbations that evaluated the predictive capacity of several empirical and first principles models at both mid- and high-latitudes in order to choose an operative model. In this work, we use three different deep learning models-a feed-forward neural network, a long short-term memory recurrent network and a convolutional neural network-to forecast the horizontal component of the ground magnetic field rate of change (dB H /dt) over 6 different ground magnetometer stations and to compare as directly as possible with the original GEM challenge. We find that, in general, the models are able to perform at similar levels to those obtained in the original challenge, although the performance depends heavily on the particular storm being evaluated. We then discuss the limitations of such a comparison on the basis that the original challenge was not designed with machine learning algorithms in mind.

Victor A. Pinto↗

Detecting Process Equipment Failures Using Acoustic Data and Machine Learning

Nuclear power plant (NPP) process equipment such as fans, motors, valves, and pumps generate frequent or continuous noise, and deviations from the normal operational sounds made by this equipment can indicate potential issues. These deviations can be identified via automated acoustic anomaly detection, which involves using acoustic sensors (i.e., microphones) alongside detection algorithms to continuously monitor for changes in acoustic signatures. This task is made challenging by the substantial background noise that exists, such as operators opening and closing doors, manipulating valves, and conversing—in addition to typical plant noises. In collaboration with a nuclear power utility partner, this effort assessed the efficacy of acoustic anomaly detection when using a specific acoustic sensor that compresses data into a fixed set of features that are transferable over a standard Internet of Things communication protocol, thereby improving usability but potentially degrading detection performance. Two methods of performing automated acoustic anomaly detection were evaluated: one-class support vector machine (OC-SVM) and isolation forest (iForest). To enable the use of high-quality acoustic data encompassing both normal and anomalous conditions, the study utilized the publicly available Malfunctioning Industrial Machine Investigation and Inspection dataset, which includes real measured acoustic sensor data for a range of equipment types, model numbers, and signal-to-noise ratios (SNRs), along with a benchmark set of detection results. Using this dataset, the methods were tested and then compared against the benchmark results. The results indicated that although the specific acoustic sensor did not enable as rich a feature set extraction, the proposed methods with the limited feature set performed just as well. This provides solid justification for both the methods and the use of the proposed acoustic sensor.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Utilizing Machine Learning to Improve Neutralization Potency of an HIV-1 Antibody Targeting the gp41 N-Heptad Repeat

The N-heptad repeat (NHR) of the HIV-1 gp41 prehairpin intermediate (PHI) is an attractive potential vaccine target with high sequence conservation across diverse strains. However, despite the potency of NHR-targeting peptides and clinical efficacy of the NHR-targeting entry inhibitor enfuvirtide, no potently neutralizing NHR-directed monoclonal antibodies (mAbs) nor antisera have been identified or elicited to date. The lack of potent NHR-binding mAbs both dampens enthusiasm for vaccine development efforts at this target and presents a barrier to performing passive immunization experiments with NHR-targeting antibodies. To address this challenge, we previously developed an improved variant of the NHR-directed mAb D5, called D5_AR, which is capable of neutralizing diverse tier-2 viruses. Building on that work, here we present the 2.7Å-crystal structure of D5_AR bound to NHR mimetic peptide IQN17. We then utilize protein language models and supervised machine learning to generate small (n < 100) libraries of D5_AR variants that are subsequently screened for improved neutralization potency. We identify a variant with 5-fold improved neutralization potency, D5_FI, which is the most potent NHR-directed monoclonal antibody characterized to date and exhibits broad neutralization of tier-2 and −3 pseudoviruses as well as replicating R5 and X4 challenge strains. Additionally, our work highlights the ability of protein language models to efficiently identify improved mAb variants from relatively small libraries.

Biopolymers↗

Anomaly Detection in Electronic Health Records Across Hospital Networks: Integrating Machine Learning With Graph Algorithms

In a large hospital system, a network of hospitals relies on electronic health records (EHRs) to make informed decisions regarding their patients in various clinical domains. Consequently, the dependability of the health information technology (HIT) systems responsible for collecting EHR data is of utmost importance for patient safety. Recently, novel methods and tools aimed at identifying anomalies in EHR data to bolster the reliability of HIT systems have been introduced. However, these existing methods and tools primarily concentrate on individual hospitals, which limits our understanding of system-wide anomalous events and their potential impact on patient safety across multiple hospitals. In this article, we introduce a new approach to detecting anomalies in EHR data within a network of hospitals. This is achieved by combining advanced machine learning techniques with graph algorithms to create a tool capable of swiftly identifying and responding to deviations. Our proposed approach employs a combination of five machine learning models, harnessing the unique strengths of each model to provide a more robust detection system. The detected anomalies are then represented as graphs, allowing us to recognize patterns across the hospital network. This aids in identifying anomalies that span multiple medical facilities, potentially indicating broader system-level risks. Extensive real-world testing of our approach demonstrated its ability to offer actionable insights compared to existing methods. Additionally, its scalable design ensures seamless integration into existing HIT infrastructures.

Niu, Haoran [Oak Ridge National Laboratory (ORNL),↗

Evaluation of GlassNet for physics-informed machine learning of glass stability and glass-forming ability

Glassy materials form the basis of many modern applications, including nuclear waste immobilization, touch-screen displays, and optical fibers, and also hold great potential for future medical and environmental applications. However, their structural complexity and large composition space make design and optimization challenging for certain applications. Of particular importance for glass processing and design is an estimate of a given composition's glass-forming ability (GFA). However, there remain many open questions regarding the underlying physical mechanisms of glass formation, especially in oxide glasses. It is apparent that a proxy for GFA would be highly useful in glass processing and design, but identifying such a surrogate property has proven itself to be difficult. While glass stability (GS) parameters have historically been used as a GFA surrogate, recent research has demonstrated that most of these parameters are not accurate predictors of the GFA of oxide glasses. Here, in this work, we explore the application of an open-source pre-trained neural network model, GlassNet, that can predict the characteristic temperatures necessary to compute GS with reasonable performance and assess the feasibility of using these physics-informed machine learning (PIML)-predicted GS parameters to estimate GFA. In doing so, we track the uncertainties at each step of the computation—from the original ML prediction errors to the compounding of errors during GS estimation, and finally to the final estimation of GFA. While GlassNet exhibits reasonable accuracy on all individual properties, we observe a large compounding of error in the combination of these individual predictions for the PIML prediction of GS, finding that random forest models offer similar accuracy to GlassNet. We also break down the performance of GlassNet on different glass families and find that the error in GS prediction is correlated with the error in crystallization peak temperature prediction. Lastly, we utilize this finding to assess the relationship between top-performing GS parameters and GFA for two ternary glass systems: sodium borosilicate and sodium iron phosphate glasses. We conclude that to obtain true ML predictive capability of GFA, significantly more data needs to be collected.

36 MATERIALS SCIENCE↗

Avoiding Braess' Paradox Through Collective Intelligence

In an Ideal Shortest Path Algorithm (ISPA), at each moment each router in a network sends all of its traffic down the path that will incur the lowest cost to that traffic. In the limit of an infinitesimally small amount of traffic for a particular router, its routing that traffic via an ISPA is optimal, as far as cost incurred by that traffic is concerned. We demonstrate though that in many cases, due to the side-effects of one router's actions on another routers performance, having routers use ISPA's is suboptimal as far as global aggregate cost is concerned, even when only used to route infinitesimally small amounts of traffic. As a particular example of this we present an instance of Braess' paradox for ISPA'S, in which adding new links to a network decreases overall throughput. We also demonstrate that load-balancing, in which the routing decisions are made to optimize the global cost incurred by all traffic currently being routed, is suboptimal as far as global cost averaged across time is concerned. This is also due to "side-effects", in this case of current routing decision on future traffic. The theory of COllective INtelligence (COIN) is concerned precisely with the issue of avoiding such deleterious side-effects. We present key concepts from that theory and use them to derive an idealized algorithm whose performance is better than that of the ISPA, even in the infinitesimal limit. We present experiments verifying this, and also showing that a machine-learning-based version of this COIN algorithm in which costs are only imprecisely estimated (a version potentially applicable in the real world) also outperforms the ISPA, despite having access to less information than does the ISPA. In particular, this COIN algorithm avoids Braess' paradox.

Wolpert , David H.↗

Fast Assessment of Metal Performance through Dislocation Physics and Machine Learning

The microstructure of metals is key to their mechanical properties. The types, density, composition and morphology of crystal defects all have pronounced impact on the properties. Changes to the microstructure occurring during processing and use can be very striking. The emerging technology additive manufacturing (AM) has the potential to improve performance by allowing optimized designs, but the process and environments can lead to unusual microscale features whose properties must be understood and characterized to enable higher technological readiness levels and application. Experimentally, an extensive evaluation of mechanical properties of 3D printed metals is a challenge, and anomalous effects related to the AM process add complexity. We present a new machine learning (ML) model predicting mechanical response based on dislocation mediated plasticity simulations. A large set of 3D discrete dislocation dynamics simulations with wide ranges of loading conditions is transformed to preprocessed data ready for training with the ML model. The trained model can predict the mechanical response of Mo30W for a given microstructure evolution, providing key information essential for optimization of AM processing.

Jaehyun Cho↗

Fast Assessment of Metal Performance through Dislocation Physics and Machine Learning

The microstructure of metals is key to their mechanical properties. The types, density, composition and morphology of crystal defects all have pronounced impact on the properties. Changes to the microstructure occurring during processing and use can be very striking. The emerging technology additive manufacturing (AM) has the potential to improve performance by allowing optimized designs, but the process and environments can lead to unusual microscale features whose properties must be understood and characterized to enable higher technological readiness levels and application. Experimentally, an extensive evaluation of mechanical properties of 3D printed metals is a challenge, and anomalous effects related to the AM process add complexity. We present a new machine learning (ML) model predicting mechanical response based on dislocation mediated plasticity simulations. A large set of 3D discrete dislocation dynamics simulations with wide ranges of loading conditions is transformed to preprocessed data ready for training with the ML model. The trained model can predict the mechanical response of Mo30W for a given microstructure evolution, providing key information essential for optimization of AM processing.

Jaehyun Cho↗

Mitigating Impact Through Community-Engaged Flood Modeling

Urban pluvial flooding poses a growing threat to the city of Baltimore, driven by heavy rainfall, increased impervious area, and aging infrastructure. Adapting to the risks posed by pluvial flooding is critical for building greater climate resiliency in Baltimore's Inner Harbor Watershed. This study addresses these challenges through community-informed decision analysis, which uses hydrologic modeling and optimization tools to identify robust flooding adaptation pathways. We will collaborate with community partners to identify key concerns and objectives regarding flooding. These concerns have been purposefully built in to a combined surface-subsurface dynamic flow simulation model. Model outputs are used to identify flooding locations within the Inner Harbor, and to test adaptation methods. Machine learning will be used search for solutions which meet diverse environmental, financial, and social goals, and solution performance will be examined under a wide range of potential future climatic conditions and integrated with an adaptive planning approach. This novel set of adaptation pathways will enhance the City's capacity to respond to evolving pluvial flood risk.

climate resilience↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Spatiotemporal Dynamics of the Relative Abundance of Soil Nutrient‐Degrading Enzyme‐Encoding Genes Across Continental US Ecoregions

Understanding the spatiotemporal patterns in the relative abundance of soil extracellular enzyme‐encoding genes is critical for predicting microbial responses to environmental change and their potential role in nutrient cycling. Yet, integrating novel metagenomic observations with spatiotemporal environmental gradients to infer regional patterns and future trajectories has remained unclear. To address this gap, we applied a machine learning (ML) approach, integrating soil metagenomic data with environmental variables—soil properties, topography, vegetation, and climate—to predict the relative abundance of enzyme‐encoding genes for soil carbon (C), nitrogen (N), and phosphorus (P) across surface soils of the continental United States. We assessed potential responses under future emission scenarios (SSP2‐4.5 and SSP5‐8.5) by comparing a baseline (1985–2014) to a future period (2071–2100). The ML model explained 57%–63% of baseline variation. Precipitation was identified as the most influential factor for the relative abundance of C‐ and N‐degrading enzyme‐encoding genes, while slope length, representing horizontal distance that water can travel downslope, was the primary driver for P‐degrading enzyme‐encoding genes abundance. Projections revealed spatially heterogeneous shifts across continental US ecoregions: the relative abundance of C‐ and N‐degrading enzyme‐encoding genes decreased in wetter ecoregions and increased in drier ecoregions under future climate, while P‐degrading enzyme‐encoding genes abundance decreased significantly in semiarid and Mediterranean ecoregions. This study demonstrates the utility of metagenomic data for mapping soil genetic potential and predicting its regional response to environmental change, to inform ecosystem management strategies.

extracellular enzyme-encoding genes↗

Global Landslide Hazard Assessment for Situational Awareness (LHASA) Version 2: New Activities and Future Plans

A remote sensing-based system has been developed to characterize the potential for rainfall-triggered landslides across the globe in near real-time. The Landslide Hazard Assessment for Situational Awareness (LHASA) model uses a decision tree framework to combine a static susceptibility map derived from information on slope, rock characteristics, forest loss, distance to fault zones and distance to road networks with satellite precipitation estimates from the Global Precipitation Measurement (GPM) mission. Since 2016, the LHASA model has been providing near real-time and retrospective estimates of potential landslide activity. Results of this work are available at https://landslides.nasa.gov. In order to advance LHASA’s capabilities to characterize landslide hazards and impacts dynamically, we have implemented a new approach that leverages machine learning, new parameters, and new inventories. LHASA 2.0 uses the XGBoost machine learning model to bring in dynamic variables as well as additional static variables to better represent landslide hazard globally. Global rainfall forecasts are also being evaluated to provide a 1-3 day forecast of potential landslide activity. Additional factors such as recent seismicity and burned areas are also being considered to represent the preconditioning or changing interactions with subsequent rainfall over affected areas. A series of parameters are being tested within this structure using NASA’s Global Landslide Catalog as well as many other event-based and multi-temporal inventories mapped by the project team or provided by project partners. In addition to estimates of landslide hazard, LHASA Version 2 will incorporate dynamic estimates of exposure including population, roads and infrastructure to highlight the potential impacts that rainfall-triggered landslides. The ultimate goal of LHASA Version 2.0 is to approximate the relative probabilities of landslide hazard and exposure across different space and time scales to inform hazard assessment retrospectively over the past 20 years, in near real-time, and in the future. In addition to the hazard. This presentation will outline the new activities for LHASA Version 2.0 and present some next steps for this system.

Dalia Kirschbaum↗

Evaluation of Technology Concepts for Traffic Data Management and Relevant Audio for Datalink in Commercial Airline Flight Decks

Datalink is currently operational for departure clearances and in oceanic environments and is currently being tested in high altitude domestic enroute airspace. Interaction with even simple datalink clearances may create more workload for flight crews than the voice system they replace if not carefully designed. Datalink may also introduce additional complexity for flight crews with hundreds of uplink messages now defined for use. Finally, flight crews may lose airspace awareness and operationally relevant information that they normally pickup from Air Traffic Control (ATC) voice communications with other aircraft (i.e., “party-line” transmissions). Once again, automation may be poised to increase workload on the flight deck for incremental benefit. Datalink implementation to support future air traffic management concepts needs to be carefully considered, understanding human communication norms and especially, the change from voice- to text-based communications modality and its effect on pilot workload and situation awareness. Increasingly autonomous systems, where autonomy is designed to support human-autonomy teaming, may be suited to solve these issues. NASA is conducting research and development of increasingly autonomous systems, utilizing machine-learning algorithms seamlessly integrated with humans whereby task performance of the combined system is significantly greater than the individual components. Increasingly autonomous systems offer the potential for significantly improved levels of performance and safety that are superior to either human or automation alone. Two increasingly autonomous systems concepts - a traffic data manager and a conversational co-pilot - were developed to intelligently address the datalink issues in a complex, future state environment with significant levels of traffic. The system was tested for suitability of datalink usage for terminal airspace. The traffic data manager allowed for automated declutter of the Automatic Dependent Surveillance-Broadcast (ADS-B) display. The system determined relevant traffic for display based on machine learning algorithms trained by experienced human pilot behaviors. The conversational co-pilot provided relevant audio air traffic control messages based on context and proximity to ownship. Both systems made use of the connected aircraft concepts to provide intelligent context to determine relevancy above and beyond proximity to ownship. A human-in-the-loop test was conducted in NASA Langley Research Center’s Integration Flight Deck B-737-800 simulator to evaluate the traffic data manager and the conversational co-pilot. Twelve airline crews flew various normal and non-normal procedures and their actions and performance were recorded in response to the procedural events. This paper details the flight crew performance and evaluation during the events.

Etherington, Timothy↗

Hydropower potential derived from streamflow extremes for Alaska, USA

Alaska is an expansive region known for its abundant natural resources, including thousands of miles of streams and rivers. These rivers represent potential opportunities for future hydropower development that could provide reliable energy supply for local communities. There is limited long-term high temporal resolution streamflow data available for the region, making data-driven estimates of potential hydropower and its variability across the state challenging. This study provides a novel data-driven approach for hydropower capacity estimation across Alaska. We use supervised machine learning to develop a relationship between the daily and peak flow duration curves in order to augment the size of our dataset from 44 sites to 67 sites. We perform a stochastic hydropower estimation across the 67 sites and identify approximately 1000 MW of total potential hydropower capacity distributed across these sites. Our study provides the first step towards more comprehensive hydropower estimation for this critical region, highlighting the need for future work integrating high-resolution spatial data, community needs, and economic constraints in estimates of potential hydropower development in Alaska.

Hydropower↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

Poster presented at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. This poster highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Applying Machine Learning to Jet Noise Prediction

This presentation summarizes the application of machine learning to jet noise data in an effort to predict the resulting noise from the interaction between a jet and a hard surface. The Aero-Acoustic Propulsion Laboratory at the NASA Glenn Research Center has acquired the noise resulting from the interaction between a jet and metal plate over a range of surface placements (e.g. plate lengths and positions) and a range of jet flow configurations. For each configuration, the noise was measured at 24 observer locations via a microphone array centered around the jet nozzle. An artificial neural network developed with Keras and TensorFlow was trained on the data to predict an 88-band spectrum as a function of surface placement, jet conditions, and observer location. Analysis of the machine learning models provide insight into which experimental parameters contribute more to the noise and which parameters could potentially be removed entirely to simplify future experiments. Preliminary results will be discussed and presented via a live demonstration of the software, which outputs a sound spectrum in real-time with user-inputted jet-surface configurations.

Dowdall, Jonny↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

This is the conference paper accompanying a poster presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24 , 2024. This work highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗