Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Low-Cost Sensor Performance Intercomparison, Correction Factor Development, and 2+ Years of Ambient PM2.5 Monitoring in Accra, Ghana

Particulate matter air pollution is a leading cause of global mortality, particularly in Asia and Africa. Addressing the high and wide-ranging air pollution levels requires ambient monitoring, but many low- and middle-income countries (LMICs) remain scarcely monitored. To address these data gaps, recent studies have utilized low-cost sensors. These sensors have varied performance, and little literature exists about sensor intercomparison in Africa. By colocating 2 QuantAQ Modulair-PM, 2 PurpleAir PA-II SD, and 16 Clarity Node-S Generation II monitors with a reference-grade Teledyne monitor in Accra, Ghana, we present the first intercomparisons of different brands of low-cost sensors in Africa, demonstrating that each type of low-cost sensor PM2.5 is strongly correlated with reference PM2.5, but biased high for ambient mixture of sources found in Accra. When compared to a reference monitor, the QuantAQ Modulair-PM has the lowest mean absolute error at 3.04 μg/m3, followed by PurpleAir PA-II (4.54 μg/m3) and Clarity Node-S (13.68 μg/m3). We also compare the usage of 4 statistical or machine learning models (Multiple Linear Regression, Random Forest, Gaussian Mixture Regression, and XGBoost) to correct low-cost sensors data, and find that XGBoost performs the best in testing (R2: 0.97, 0.94, 0.96; mean absolute error: 0.56, 0.80, and 0.68 μg/m3 for PurpleAir PA-II, Clarity Node-S, and Modulair-PM, respectively), but tree-based models do not perform well when correcting data outside the range of the colocation training. Therefore, we used Gaussian Mixture Regression to correct data from the network of 17 Clarity Node-S monitors deployed around Accra, Ghana, from 2018 to 2021. We find that the network daily average PM2.5 concentration in Accra is 23.4 μg/m3, which is 1.6 times the World Health Organization Daily PM2.5 guideline of 15 μg/m3. While this level is lower than those seen in some larger African cities (such as Kinshasa, Democratic Republic of the Congo), mitigation strategies should be developed soon to prevent further impairment to air quality as Accra, and Ghana as a whole, rapidly grow.

Humidity↗

Detecting Risk and Anomalies in Airplane Dynamics Through Entropic Analysis of Time Series Data

Despite recent efforts to move away from traditional threshold exceedance detection methods for aircraft state monitoring, modern aircraft still rely on safety thresholds to communicate to pilots the identification of an anomaly in the aircraft when a threshold is surpassed. Current anomaly detection methods mainly depend on uninterpretable machine learning models to learn complex patterns and relationships contained in the time series data of aircraft. Although these methods are capable of identifying known anomalies, their deficiency in interpretability presents a challenge when translating them to different aircraft. To overcome this deficiency, entropic analysis of aircraft dynamics seeks to characterize the complexity, or lack thereof, of the aircraft dynamics prior to the development of a risk scenario. This complexity characterization provides a more straightforward summary of state changes in the dynamics of flight variables. To build a foundation for entropic analysis, we analyzed the complexity of unstable approaches, an anomalous event present in many of today’s aviation accidents. The analysis revealed a statistically significant difference in the complexity distribution of flight variables under a stable approach versus an unstable approach. These differences in complexity were especially notable minutes before an approach was identified as unstable. Moreover, the multiscale entropic analysis revealed the presence of signal complexity at multiple time scales across multiple time windows before landing. By capturing state changes and corrections in the aircraft dynamics using entropy, advanced, yet still interpretable, sensor systems based on entropic frameworks from this study can be constructed in the future using classical machine learning approaches.

Risk detection↗

Algorithmic Detection of Elemental Biosignatures

Machine learning models that classify a sample as indicative or non-indicative of life could play an important role in life-detection missions. Their predictions result from agnostic algorithms and thereby add redundancy to judgements resulting from human expertise. Additionally, their important features can reveal the most informative measurements within the operational constraints of a life-detection mission. The Ladder of Life Detection (Neveu 2018) identifies the need for an understanding of how combinations of multiple biosignatures affect overall confidence. The present work provides a starting point to answer this need, and future work will expand the data types to obtain even more predictive combinations of features. Elemental abundance was chosen as a starting set of features due to its availability in diverse sample types, which are needed to train a generalizable model. A standardized dataset was collected, including 35 non-indicative, e.g., lunar rock, basalt; 19 indicative mixed, e.g., seawater, agricultural soil; 46 indicative non-alive, e.g., coal, chalk; and 10 indicative alive, e.g., biofilm, bacteria. This dataset could be valuable for complementary biosignature research. The samples were standardized to the same limit of detection of a simulated mission scenario. Four classification models were used: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), and Gaussian naïve Bayes (GNB). To obtain feature importances, KNN was run on three principal components of the training data and LR and SVM were run with L1 and L2 regularization. The performances and feature importances of the six model variants on 40:60 train to validation ratios were assessed with Monte Carlo simulations. ROC AUC and mean accuracy scores ranged between 82% - 94%, with sensitivity greater than specificity. For indicative of life predictors, all models had C and Ca as strong and Cl as medium; a majority of models had N, K, and P as medium. For non-indicative of life predictors, all models had Si as strong, and a majority of models had Mg, Al, and Ti as medium. Varied elements were Fe (slightly non-indicative), H (slightly indicative), O (widely varied), Na, Mn, and S. These results serve as a proof of concept and suggest important elemental signals beyond merely the CHNOPS of Earth-based life.

Algorithmic↗

Predicting Two-Dimensional Airfoil Performance Using Graph Neural Networks

Computer simulations require the use of meshes to simulate geometries. These meshes capture important geometric features of the design and can be used in machine learning modeling. This report explores the use of graph neural networks (GNNs) to learn features from two-dimensional (2D) airfoil designs represented as a set of nodes connected using edges. This type of network is common in aerospace applications: most geometries are represented as a mesh in order to perform analysis. The objective of this work is to use GNNs to predict the performance of 2D airfoils generated using the program XFOIL. The predicted performance parameters include bulk quantities such as coefficients of lift (C L ), drag (C d , C dp ), moment (C m ), and node-specific quantities such as coefficient of pressure (C p ). In this report, a spline convolutional graph-based neural network is compared with deep learning neural networks to predict both bulk and node-specific quantities. The findings indicate the GNNs are able to predict bulk quantities quite well; however, when the number of outputs is increased, the deep neural network (DNN) proves to be better in its prediction capability. Two different normalization strategies were compared in the training of both GNNs and DNNs: minmax and standard deviation. In both types of networks, standard deviation scaling proved to be the best.

machine learning↗

Using Historical Data to Automatically Identify Air-Traffic Control Behavior

This project seeks to develop statistical-based machine learning models to characterize the types of errors present when using current systems to predict future aircraft states. These models will be data-driven - based on large quantities of historical data. Once these models are developed, they will be used to infer situations in the historical data where an air-traffic controller intervened on an aircraft's route, even when there is no direct recording of this action.

trajectory generation↗

Data Mining for Understanding and Impriving Decision-Making Affecting Ground Delay Programs

The continuous growth in the demand for air transportation results in an imbalance between airspace capacity and traffic demand. The airspace capacity of a region depends on the ability of the system to maintain safe separation between aircraft in the region. In addition to growing demand, the airspace capacity is severely limited by convective weather. During such conditions, traffic managers at the FAA's Air Traffic Control System Command Center (ATCSCC) and dispatchers at various Airlines' Operations Center (AOC) collaborate to mitigate the demand-capacity imbalance caused by weather. The end result is the implementation of a set of Traffic Flow Management (TFM) initiatives such as ground delay programs, reroute advisories, flow metering, and ground stops. Data Mining is the automated process of analyzing large sets of data and then extracting patterns in the data. Data mining tools are capable of predicting behaviors and future trends, allowing an organization to benefit from past experience in making knowledge-driven decisions. The work reported in this paper is focused on ground delay programs. Data mining algorithms have the potential to develop associations between weather patterns and the corresponding ground delay program responses. If successful, they can be used to improve and standardize TFM decision resulting in better predictability of traffic flows on days with reliable weather forecasts. The approach here seeks to develop a set of data mining and machine learning models and apply them to historical archives of weather observations and forecasts and TFM initiatives to determine the extent to which the theory can predict and explain the observed traffic flow behaviors.

data mining↗

Probabilistic Prognosis of Non-Planar Fatigue Crack Growth

Quantifying the uncertainty in model parameters for the purpose of damage prognosis can be accomplished utilizing Bayesian inference and damage diagnosis data from sources such as non-destructive evaluation or structural health monitoring. The number of samples required to solve the Bayesian inverse problem through common sampling techniques (e.g., Markov chain Monte Carlo) renders high-fidelity finite element-based damage growth models unusable due to prohibitive computation times. However, these types of models are often the only option when attempting to model complex damage growth in real-world structures. Here, a recently developed high-fidelity crack growth model is used which, when compared to finite element-based modeling, has demonstrated reductions in computation times of three orders of magnitude through the use of surrogate models and machine learning. The model is flexible in that only the expensive computation of the crack driving forces is replaced by the surrogate models, leaving the remaining parameters accessible for uncertainty quantification. A probabilistic prognosis framework incorporating this model is developed and demonstrated for non-planar crack growth in a modified, edge-notched, aluminum tensile specimen. Predictions of remaining useful life are made over time for five updates of the damage diagnosis data, and prognostic metrics are utilized to evaluate the performance of the prognostic framework. Challenges specific to the probabilistic prognosis of non-planar fatigue crack growth are highlighted and discussed in the context of the experimental results.

Leser, Patrick E.↗

Applying Machine Learning to Jet Noise Prediction

This presentation summarizes the application of machine learning to jet noise data in an effort to predict the resulting noise from the interaction between a jet and a hard surface. The Aero-Acoustic Propulsion Laboratory at the NASA Glenn Research Center has acquired the noise resulting from the interaction between a jet and metal plate over a range of surface placements (e.g. plate lengths and positions) and a range of jet flow configurations. For each configuration, the noise was measured at 24 observer locations via a microphone array centered around the jet nozzle. An artificial neural network developed with Keras and TensorFlow was trained on the data to predict an 88-band spectrum as a function of surface placement, jet conditions, and observer location. Analysis of the machine learning models provide insight into which experimental parameters contribute more to the noise and which parameters could potentially be removed entirely to simplify future experiments. Preliminary results will be discussed and presented via a live demonstration of the software, which outputs a sound spectrum in real-time with user-inputted jet-surface configurations.

Dowdall, Jonny↗

Reusing Data and Metadata to Create New Metadata Through Machine-Learning & Other Programmatic Methods

Recent improvements in natural language processing (NLP) enable metadata to be created programmatically from reused original metadata or even the dataset itself. Transfer-learning applied to NLP has greatly improved performance and reduced training data requirements. In this talk, we’ll compare machine-generated metadata to human-generated metadata and discuss characteristics of metadata and data archives that affect suitability for machine-learning reuse of metadata. Where as human-generated metadata is often populated once, populated from the perspective of data supplier, populated by many individuals with different words for the same thing, and limited in length, machine-generated metadata can be updated any number of times, generated from the perspective of any user, constrained to a standardized set of terms that can be evolved over time, and be any length required. Machine-learning generated metadata offers benefits but also additional needs in terms of version control, process transparency, human-computer interaction, and IT requirements. As a successful example, we’ll discuss how a dataset of abstracts and associated human-tagged keywords from a standardized list of several thousand keywords were used to create a machine-learning model that predicted keyword metadata for open-source code projects on code.nasa.gov. We’ll also discuss a less successful example from data.nasa.gov to show how data archive architecture and characteristics of initial metadata can be strong controls on how easy it is to leverage programmatic methods to reuse metadata to create additional metadata.

Gosses, Justin↗

Assessing the Use of SAR/Optical Data Fusion and TensorFlow for Improved Mangrove Mapping

Mangrove forests are found in intertidal zones of tropical regions around the world and provide important ecological and economic benefits – they are considered carbon sequesters, habitats for flora and fauna, and natural barriers to hurricanes and tsunamis. Wood from mangrove forests are used as fuel and building materials in surrounding coastal communities, therefore promoting local livelihoods. Despite the importance of these ecosystems, mangrove forests have historically been degraded in natural processes such as severe weather, and anthropogenic factors like conversion to agriculture and aquaculture. This study assesses change in mangrove forests in Nigeria and Mozambique from 2015 to 2018 using SAR and optical data fusion. Due to frequent cloud cover over the study area, SAR and optical data is fused to obtain gap-free imagery without clouds. Landsat-8 OLI and Sentinel-1 imagery is fused with TensorFlow, an open source platform used in developing machine learning models. The resulting images are classified to discriminate mangrove forest cover from other land cover types, and change is estimated using image differencing. Understanding the rates and magnitude of mangrove change across space and time can aid in identifying priority areas for forest regeneration, and can help construct sustainable management practices for the future.

Strattman, Katherine↗

TPSAS-NF1676L-32345-DND

Interest in the use of Raman spectrometers has seen an increase in the fields of geology and planetary sciences due to the non-destructive insight Raman spectra may provide into the molecular makeup of a given sample. Advancements in Raman spectrometer hardware have allowed for compact instruments to have deployment capabilities directly on interplanetary missions, flexible usage conditions requiring no sample collection/preparation, and no need for daylight radiation shielding. As the amount of science which can be collected from a Raman spectrometer in a given amount of time increases, a bottleneck will be created in data analysis which leaves a need for a faster method of spectral data classification. Recent studies have shown that machine learning models are able to solve this problem by achieving high-accuracy classification. Liu et al4 found the convolutional neural network (CNN) held the highest classification accuracy (96% top 5) for single sample Raman data.

A Atkinson↗

A Machine Learning-Based Cloud Detection and Thermodynamic Phase Classification Algorithm using Passive Spectral Observations

We trained two Random Forest (RF) machine-learning models for cloud mask and cloud thermodynamic phase detection using spectral observations from VIIRS on Suomi NPP (SNPP). Observations from CALIOP were carefully selected to provide reference labels. The two RF models were trained for all-day and daytime-only conditions using a 4-year collocated VIIRS/CALIOP dataset from 2013 to 2016. Due to the orbit difference, the collocated CALIOP and SNPP VIIRS training samples cover a broad viewing zenith angle range, which is a great benefit to overall model performance. The all-day model uses 3 VIIRS infrared (IR) bands (8.6,11, and 12 μm) and the daytime model uses 5 Near-IR (NIR) and Shortwave-IR (SWIR) bands (0.86, 1.24, 1.38, 1.64 and 2.25 μm) together with the 3 IR bands to detect clear, liquid water, and ice cloud pixels. Up to 7 surface types, namely, ocean/water, forest, cropland, grassland, snow/ice, barren/desert, and shrubland, were considered separately to enhance performance for both models. Detection of cloudy pixels and thermodynamic phase with the two RF models were compared against collocated CALIOP products from 2017. It is shown that, with a conservative screening process that excludes the most challenging cloudy pixels for passive remote sensing, the two RF models have high accuracy rates in comparison with the CALIOP reference for both cloud detection and thermodynamic phase. Other existing SNPP VIIRS and Aqua MODIS cloud mask and phase products are also evaluated, with results showing that the two RF models and the MODIS MYD06 optical property phase product are the top 3 algorithms with respect to lidar observations during the daytime. During the nighttime, the RF all-day model works best for both cloud detection and phase, in particular for pixels over snow/ice surfaces. The present RF models can be extended to other similar passive instruments if training samples can be collected from CALIOP or other lidars. However, the quality of reference labels and potential sampling issues that may impact model performance would need further attention.

cloud detection↗

Measurement of material recession and shock standoff in plasma windtunnel using neural nets

Arcjets are plasma wind tunnels used to test the performance of heatshield materials for spacecraft atmospheric entry. These facilities present an extremely harsh flow environment with heat fluxes up to 10^9 W/m^2 for up to 30 minutes. The plasma is low-temperature (~1 eV) but high pressure (> 10 kPa) creating high-enthalpy supersonic flows similar to atmospheric entry conditions. Typically, material samples are measured before and after a test to characterize the total recession. However, this does not capture time-dependent effects such as material expansion and non-linear recession. This work will present new analysis of arcjet test videos which measure both the time-dependent 2D recession of the material samples and the shock standoff distance. New results showing non-linear material erosion rates will be highlighted. The material and shock edges are extracted from the videos by training and applying a convolutional neural network. Due to the consistent camera settings, the machine learning model achieves high accuracy (~99%) on new data with only a small number of training frames (~80). The new results will be discussed in the context of temperature dependent plasma-surface interaction.

Magnus A Haw↗

Measurement of Material Recession and Shock Standoff in Plasma Windtunnel using Neural Nets

Arcjets are plasma wind tunnels used to test the performance of heatshield materials for spacecraft atmospheric entry. These facilities present an extremely harsh flow environment with heat fluxes up to 109 W/m2 for up to 30 minutes. The plasma is low-temperature (∼1 eV) but high pressure (> 10 kPa) creating high-enthalpy supersonic flows similar to atmospheric entry conditions. Typically, material samples are measured before and after a test to characterize the total recession. However, this does not capture time-dependent effects such as material expansion and non-linear recession. This work will present new analysis of arcjet test videos which measure both the time-dependent 2D recession of the material samples and the shock standoff distance. The results show non-linear time-dependent effects are present for some conditions. The material and shock edges are extracted from the videos by training and applying a convolutional neural network. Due to the consistent camera settings, the machine learning model achieves high accuracy (± 2 px) relative to manually segmented images with only a small number of training frames (80).

Neural network↗

A materials-informatics based study of solid electrolytes and protective coatings for Li batteries

All-solid-state batteries with Li metal anode can address the safety issues surrounding traditional Li-ion batteries as well as the demand for higher energy densities. However, the development of solid electrolytes and protective coatings simultaneously possessing high ionic conductivity and wide electrochemical stability has proven to be a challenge. Here, we present a data-driven approach to explore the Li compound space for promising solid electrolytes and coatings. This is accomplished through the generation of a large database of battery-related materials properties of Li compounds by computing Li+ migration barriers using bond-valence-based pair potentials, and stability windows using density functional theory energies. Using this database, we implement machine learning models that can accurately predict migration barriers and electrochemical stability windows for any new Li compound. Through feature engineering, we ensure that our models are both accurate and interpretable. We perform feature importance analysis on our models to highlight materials properties that can be tuned for future design of coatings/electrolytes. Our database and informatics approach provide a valuable tool for the rapid discovery of new solid-state battery chemistries.

Solid state batteries↗

Highland Lakes Water Resources: Using NASA Earth Observations to Improve Detection Systems for Harmful Algal Events in the Highland Lakes in Central Texas

Beginning in 2019, harmful algal events in Austin, Texas, caused canine deaths in the Lady Bird Lake and Lake Travis reservoirs. These reservoirs are part of the larger Highland Lakes chain, managed by the Lower Colorado River Authority (LCRA) and the City of Austin Department of Watershed Protection (CoA DWP), which fulfill municipal, commercial, and agricultural water demands. Given the recent increase in favorable algal event conditions in central Texas, the LCRA and CoA DWP partnered with NASA DEVELOP to improve algal event early-warning systems through the application of remote sensing and machine learning. An Earth observation-based algal monitoring system will assist the responsible agencies in predicting algal conditions and communicating hazards to the public. The NASA DEVELOP team utilized Landsat 8 Operational Land Imager (OLI) and Sentinel-2 Multispectral Instrument (MSI) data to produce products including chlorophyll-a concentrations, cyanobacteria detections, turbidity, and water surface temperature. Chlorophyll-a concentrations were retrieved with a pre-trained machine learning model (mixture density network) and spectral indices, while the other products were derived from spectral indices. In situ field data were used to validate and quantify uncertainties for each product. The validations show strong correlations for chlorophyll-a and water surface temperature. Time series analyses of chlorophyll-a concentrations show peaks in the severe drought years (2015 and 2016). This project's resulting products enable monitoring of environmental proxies relevant to algal event presence in the Highland Lakes chain and will ultimately support water management, decision making, and risk communication.

Kaitlynn Hietpas↗

Predicting the Functional State of Protein Kinases Using Interpretable Graph Neural Networks

Kinases are a family of proteins that function as molecular switches, regulating several essential cellular activities such as cell proliferation. Dysfunctional kinases are implicated in several types of cancers and hence they are actively pursued as drug targets. Given the vast number of complex kinase structures that are available in the protein data bank (PDB), there is a necessity to develop methodologies that can identify structurally important moieties of the kinases in an automated fashion, for such techniques can be instrumental in identifying novel drug targets. In this work, we develop a graph neural network (GNN) based deep learning framework for classifying the functionally active and inactive states of a large set of eukaryotic protein kinases, making use of their 3D structure from the PDB. We show that GNN based machine learning models can classify protein states with an accuracy greater than 97%. We further use the GNN models to automatically identify regions of the kinases that are important for its function. For this purpose, Gradient-weighted Class Activation Mapping (Grad-CAM) was implemented on the protein graphs. Remarkably, Grad-CAM consistently identifies the highly conserved DFG motif as the most important part of the protein across the entire kinome, without any prior input. Other regions of the hydrophobic core such as the HRD motif were also identified by the interpretable GNN framework, consistent with the literature. We discuss the significance of each of these regions in detail.

Ashwin Ravichandran↗

Data-Driven Study of Shape Memory Behavior of Multi-Component Ni-Ti Alloys

Ni-Ti based shape memory alloys (SMAs) have found wide-spread use in aerospace, automotive, biomedical, and commercial applications owing to their favorable properties and ease of operation. Especially important for many NASA applications is the ability to tune the martensitic transformation temperature of Ni-Ti alloys by varying the alloy composition and processing conditions. Recently, researchers at NASA have compiled an extensive database of shape memory properties of materials, including over 8,000 multi-component Ni-Ti alloys containing 37 different alloying elements. Using this dataset, machine learning models are trained to predict transformation temperatures, hysteresis, and transformation strain with extremely low mean absolute errors. These models are used to learn relationships between shape memory behavior and input parameters in the composition and processing space. ML predictions are validated through new experiments. The combination of an extensive experimental dataset and accurate learning models, together, make our approach highly suitable for the rapid discovery and design of novel SMAs with targeted properties. We are not aware of any current approaches capable of predicting SMA transformation behavior over such a wide range of compositions and processing conditions.

Shape memory alloys↗