Search NASA⌕ Search

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Implementing Geometric Surface Imperfections into Sandwich Composite Cylinder Finite Element Method Models

The buckling responses of certain cylindrical shell structures are extremely sensitive to geometric surface imperfections. The NASA Engineering and Safety Center (NESC) Shell Buckling Knockdown Factor Project (SBKF) is conducting research to develop analysis-based buckling design recommendations. Experiments are used to verify the analysis-based factors, but the sensitivity of the test articles to geometric imperfections requires implementing as-manufactured imperfections into high-fidelity finite element method models. Data collection methods such as structured light scanning are used for all geometric surface data used in this work. Common preprocessing and visualization steps used in SBKF are discussed, and steps on how surface scans are prepared for implementation into a finite element model is described. The Python Tool for Implementing Geometric Imperfections in Reduced Structures (Py_TIGIRS), written specifically for the use with SBKF, is briefly described and uses eight functions to extract, modify, and write geometric imperfections into Abaqus input files. Results of the pre-processing methods and results from Py_TIGIRS are provided and compared for Composite Test Article (CTA) 8.2B. Excellent agreement between the visualized scan data and the FEM-extracted geometry is demonstrated. A brief example of why geometric surface imperfections are significant in nonlinear numerical analyses for thin cylinders in axial compression is provided as motivation to use tools such as Py_TIGIRS. Future development of Py_TIGIRS including expansion to structures of arbitrary geometry is planned.

Sandwich structures↗

Entwine Point Tiles for 3D Visualization and Querying of ICESat-2

Point Cloud data from non-optical sensors present challenges in scientific computing in both volume of data and files, even for cloud services environments. As part of the Multi-Mission Algorithm and Analysis Platform (MAAP), a joint open science platform for global biomass modelling, we’ve developed a cloud optimized workflow for using ATL08 (ICESat-2) data as a point cloud. For MAAP, the ATL08 data product is published as Entwine Point Tiles (EPT), allowing users to visualize and query the full extent of this collection interactively without pre-downloading, or preprocessing. The EPT format is a cloud-optimized point cloud data format which re-organizes points into a cloud friendly spatially indexed data structure. MAAP uses AWS S3 to store these point clouds and serves them over OGC specified APIs, 3DTiles for visualization, and WFS for querying. This workflow allows for interactive 3D visualizations in a web browser, including notebook environments and facilitates on the fly subsetting for interactive data exploration, all of which can be applied to other similar sensors.

Alex Mandel↗

A Quantitative Analysis on the Use of Supervised Machine Learning in Earth Science

Recent review papers (Ball et al., 2017; Reichstein et al., 2019) have investigated the opportunities and challenges in applying supervised machine learning (ML) techniques to Earth science problems. A common challenge is the lack of training (or labeled) data. Supervised ML, and especially deep learning (DL), require large training datasets. While there are large, open access Earth science archives, the data typically require preprocessing in preparation for supervised ML, frequently including manual labeling. Our objective is to understand the landscape of supervised ML in the Earth sciences, including which research communities have most rapidly adopted supervised ML, which algorithms are applied, and what data are used to train these algorithms. We conducted a literature survey of Earth science papers published during the last 10 years in journals from the American Geophysical Union (AGU), American Meteorological Society (AMS), the Institute of Electrical and Electronics Engineers(IEEE), and the Society of Photo-Optical Instrumentation Engineers (SPIE). We identified papers containing the terms ML, DL, or the names of individual supervised ML algorithms. "Earth science" is an additional required search term for IEEE and SPIE. We investigate trends in supervised ML usage during the 10-year study period, and manually analyzed AGU papers from 2018-2019 to enable deep-dive statistics.

Katrina S Virts↗

Analysis Ready Satellite Data

Analyais-Ready Data (ARD) specifications have gained rcent prominence in the field of land-related Earth Observations. The ARD label enables users to recognize data that need a minimum of preprocessing before analysis. However, the regular geolocation requirements make Level 2 data in other disciplines problematic. While Level 3 gridded data can satisfy the geolocation requirement, they often sacrifice spatial resolution and other information, such as extreme values. This talk outlines this dilemma with some potential approaches to it.

Analysis-Ready Data↗

Mine Discrimination Using Multispectral Imagery With Feedforward Neural Networks

Simulated mine detection was performed on a polarimetric hyperspectral imaging dataset collected by using an acousto-optic tunable filter camera. A feedforward artificial neural network was programmed to recognize predefined spectral "templates." The simulation results are provided along with the preprocessing steps and window sizes leading to mine detection without false alarms.

polarimetric↗

Application of ML/AI for Identifying Earth Science Datasets in Research Publications

NASA Data Active Archive Centers, or DAACs, ingest, store and distribute data acquired from satellites, ground systems as well as modelling data. These data are organized by the datasets, each presenting collection of files usually associated with the certain mission, instrument, processing level, parameter(s), algorithm and/or model. The number of datasets offered by a single DAAC to the public varies. GES DISC, for example, currently offers for public use approximately ~1,300 datasets. While each publicly offered dataset comes with supporting documentation, it is challenging for novice and even experienced scientists to navigate among the datasets that offer similar parameters to find the datasets for their particular research application. Supplying dataset documentation with the scientific paper citations that refer to that dataset provides means for the dataset users to educate themselves with the application research that dataset is being used in. Collecting citations of the papers that use the datasets for their research yield valuable insights into application areas of those datasets, information about usage of the dataset groups for specific applications and those application topics. It also gives insights into the “deep metrics” of the dataset usage, as opposed to the common metrics of the dataset usage such as number of users who downloaded the dataset files and volumes of downloaded data. Association of a certain scientific paper with the dataset(s) presents a challenge because most of the paper authors do not properly cite the datasets, datasets usually have cryptic names and Digital Object Identifiers (DOIs) that are used for dataset identification were assigned to the datasets only few years ago. Simple Google or online library search do not provide even meaningful fraction of the results when performed by the dataset name or DOI, however they provide too many results when the search is done by more broader terms such as mission and instrument names. Attempts to create an AI system capable to identify dataset in the scientific papers have already been made using neural networks classifiers on the basis of the dataset mission, instrument and variable name. This method was applied to NASA SEDAC, which has 41 datasets in total. In GES DISC there can be as many as ~100 datasets per mission/instrument with some of the datasets consisting of multiple variables so there is a need for more differentiating parameters for dataset identification in the paper. The approach we are currently investigating is creating AI classifiers that are based on multiple dataset features, or keywords, extracted from the NASA Earthdata Common Dataset Repository (CMR). The features are weighted based on how precisely they can identify a dataset. The classifier uses preprocessed paper text as input and searches for the CMR datasets whose feature sets are the closest to the feature sets contained in the paper. The challenges of dataset identification include variety of ways the paper authors describe the datasets in their papers and incomplete tagging of the CMR dataset description (DIFs).

Irina Gerasimov↗

Design and Analysis of Convolutional Neural Network for RF Signal Modulation Classification for In-Orbit Deployment

To effectively transmit data to and from satellites requires a complex and robust RF communication system. Commonly, several different types of signal modulations may be required to maximize satellite efficiency depending on a variety of unexpected channel impairments. We propose a neural network algorithm capable of learning these RF signal modulations using a supervised learning technique designed for low power, high-efficiency in-orbit deployment. The work presented demonstrates a convolutional neural network (CNN) capable of learning and recognizing a set of modulation schemes commonly used to transmit RF information. We are capable of recognizing the modulation scheme from the I and Q data channels directly, with no preprocessing or data conversion required other than breaking the incoming signal into a set of uniform normalized samples. We perform a network design and size analysis, showing that reasonably high accuracy can be obtained using networks with a relatively low number of trainable parameters. Given that a user of a system such as this may wish to receive a signal using a modulation scheme that the network has not previously learned, we demonstrate that transfer learning can learn new modulation schemes by retraining only the fully connected layers in the CNN. Thus, this type of network would excel in outer space deployment using high-efficiency transfer learning hardware. Modulation recognition can be performed through rapid feedforward computation, and the CNN training process is significantly simplified when learning new modulations is required.

CNN↗

DELTA: An Open-Source Framework to Simplify Deep Learning with Satellite Imagery

DELTA (Deep Earth Learning, Tools, and Analysis) is an open-source framework developed at NASA for deep learning on satellite imagery based on tensorflow. It helps simplify data engineering and preprocessing steps and reduces the need for a lot of the boilerplate code that needs written to make datasets palatable for machine learning. This lets data scientists focus on model development while DELTA handles the grunt work. This presentation will demonstrate DELTA’s functionality and share some examples from an active project using it for flood mapping.

Michael von Pohle↗

Classifying Agnostic Biosignatures using Raman, VNIR, and Elemental Data

How can we use our current wealth of terrestrial data, encompassing biogenic and abiogenic systems, to determine the distinguishing properties of life? SCOBI (Statistical Classification of Biosignature Information) uses machine learning techniques to algorithmically identify combinations of measurements that are “indicative of life”. A set of ~1000 observations, comprising elemental abundance, isotopic fractionation, VNIR reflectance, and (in progress) Raman spectra, have been assembled from existing literature and databases. The observations cover systems classified as “indicative alive” (e.g., cells, vegetation), “indicative non-alive” (e.g., fossils, teeth), “mixed indicative” (e.g., soil, pond water), or “non-indicative” (e.g., rocks, meteorites). VNIR data was preprocessed by linear interpolation from 400-2100 nm and smoothed with a Savitzky-Golay filter. To limit the amount of Earth-biochemistry-specific (non-agnostic) information included, the first five spectral features extracted were number of peaks, number of troughs, mean reflectance, mean peak width, and broadest peak width. To help further emphasize agnostic biosignatures, Earth-specific features such as chlorophylls have been manually flagged so that feature importance with and without them can be compared. Classifiers including k-nearest neighbors (KNN), Gaussian Naïve Bayes (GNB), logistic regression (LR), random forest (RF), and support vector machine (SVM) were implemented, as was a combination voting classifier. Performance metrics included false positive rates, false negative rates, and AUC with 50-50 test/train splits (Monte Carlo simulations). Key takeaways from this stage, prior to the inclusion of Raman spectra, are (1) the overall success rate of 0.933 AUC was most heavily influenced by the elemental abundance data; and (2) VNIR reflectance had the lowest classification performance with 0.52 AUC (58% of objects correctly classified). The next steps are to complete integration of Raman spectral data and to improve the approach to pre-processing and feature extraction for both types of spectral data, such as automated baseline removal, whole spectrum matching, and dimensionality reduction.

Biosignatures↗

DELTA: An Open-Source Framework to Simplify Machine Learning with Satellite Imagery

DELTA (Deep Earth Learning, Tools, and Analysis) is an open-source framework developed at NASA to simplify running and training machine learning (ML) models on satellite imagery. Users new to machine learning can run existing ML models on satellite imagery with minimal setup and configuration. For experienced ML users, DELTA helps simplify data engineering, preprocessing steps, and reduces the need for boilerplate code that needs written to make satellite imagery datasets palatable for machine learning. This lets data scientists focus on model development while DELTA handles the imagery manipulation. This presentation will demonstrate DELTA’s functionality and share some examples from an active project using it for flood mapping using imagery from multiple satellite sources

Michael von Pohle↗

Bhutan Water Resources III: Analyzing Forest Disturbances and Climate Data in Bhutan to Create a Tool for Assisting the Himalayan Environment Rhythm Observation and Evaluation Systems (HEROES) Project

Forest disturbances from bark beetle outbreaks are a major concern in Bhutan, known to cause extensive tree mortality to pine and spruce forests. The NASA DEVELOP team partnered with the Ugyen Wangchuck Institute of Conservation and Environmental Research (UWICER), the Bhutan Foundation, and the Karuna Foundation to assess forest changes for the districts of Bumthang and Haa from 2000 to 2018. The project used preprocessed meteorological data from the Climate Hazards Center Infrared Precipitation with Station (CHIRPS) and Famine Early Warning System Network Land Data Assimilation System (FLDAS), along with Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper plus (ETM+), and Landsat 8 Operational Land Imager (OLI) to assess apparent forest disturbance occurrences and observed climate trends. Shuttle Radar Topography Mission (SRTM) was used to resolve variations in elevation and slope for mountainous regions. Using the Google Earth Engine LandTrendr (LT) code algorithm, along with Landsat data, the team developed an app called Forest Disturbances Detection Toolbox (FDDT) to assess forest changes in Bhutan. The app includes climate variables for the focus districts, along with LT variables, which allows the end users to further examine the cause of disturbances. The team compared geocoordinates for known disturbances with LT disturbance detection products. Although additional work is needed in the future to validate the project end products from the FDDT, the tool will be provided to the project partners to aid forest management efforts in Bhutan.

Tashi Choden↗

Fast Assessment of Metal Performance through Dislocation Physics and Machine Learning

The microstructure of metals is key to their mechanical properties. The types, density, composition and morphology of crystal defects all have pronounced impact on the properties. Changes to the microstructure occurring during processing and use can be very striking. The emerging technology additive manufacturing (AM) has the potential to improve performance by allowing optimized designs, but the process and environments can lead to unusual microscale features whose properties must be understood and characterized to enable higher technological readiness levels and application. Experimentally, an extensive evaluation of mechanical properties of 3D printed metals is a challenge, and anomalous effects related to the AM process add complexity. We present a new machine learning (ML) model predicting mechanical response based on dislocation mediated plasticity simulations. A large set of 3D discrete dislocation dynamics simulations with wide ranges of loading conditions is transformed to preprocessed data ready for training with the ML model. The trained model can predict the mechanical response of Mo30W for a given microstructure evolution, providing key information essential for optimization of AM processing.

Jaehyun Cho↗

Using Earth Observations to Analyze Vegetation Phenology and Climatology in Bhutan to Identify Forest Disturbance

Changes in climate in the Himalayan region cause variability in temperature, precipitation, and phenology. It can also impact the health of coniferous forest ecosystems including increased damage due to aggressive forest pests. Forest disturbance from bark beetle is a major concern in Bhutan, sometimes causing extensive tree mortality to pine and spruce forests. Climatological trends and changes in vegetation phenology were analyzed and incorporated into a tool in Google Earth Engine that identified patches of forest disturbance in Bhutan. Preprocessed phenology and meteorological data from the Advanced Very High-Resolution Radiometer (AVHRR) and Terra and Aqua Moderate Resolution Imaging Spectroradiometer (MODIS), along with Climate Hazards Center Infrared Precipitation with Station (CHIRPS), Famine Early Warning System Network Land Data Assimilation System (FLDAS), Sentinel-2 Multispectral Instrument (MSI), and Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper (ETM)+, and Landsat 8 Operational Land Imager (OLI) were used within the tool. Changes in temperature, precipitation, and phenology were analyzed throughout Bhutan, and forest disturbance caused by bark beetle was investigated in two districts of the country

Tashi Choden↗

DELTA: An Open-Source Framework to Simplify Machine Learning with Satellite Imagery

DELTA (Deep Earth Learning, Tools, and Analysis) is an open-source framework developed at NASA to simplify running and training machine learning (ML) models on satellite imagery. Users new to machine learning can run existing ML models on satellite imagery with minimal setup and configuration. For experienced ML users, DELTA helps simplify data engineering, preprocessing steps, and reduces the need for boilerplate code that needs written to make satellite imagery datasets palatable for machine learning. This lets data scientists focus on model development while DELTA handles the imagery manipulation. This presentation will demonstrate DELTA’s functionality and share some examples from an active project using it for flood mapping using imagery from multiple satellite sources.

deep learning↗

Post-Landing Major Element Quantification Using SuperCam Laser Induced Breakdown Spectroscopy

The SuperCam instrument on the PerseveranceMars 2020 rover uses a pulsed 1064 nm laser to ablate targets at a distance and conduct laser induced breakdown spectroscopy (LIBS) by analyzing the light from the resulting plasma. SuperCam LIBS spectra are preprocessed to remove ambient light, noise, and the continuum signal present in LIBS observations. Prior to quantification, spectra are masked to remove noisier spectrometer regions andspectra are normalized to minimize signal fluctuations and effectsof target distance.In some cases, the spectra are also standardized or binned prior to quantification. To determine quantitative elemental compositionsof diverse geologic materials at Jezero crater, Mars, we use a suite of 1198 laboratory spectra of 334 well-characterized reference samples. The samples were selected to span a wide range of compositions and include typical silicate rocks, pure minerals (e.g.,silicates, sulfates, carbonates, oxides),more unusual compositions (e.g.,Mn oreand sodalite), andreplicates of the sintered SuperCam calibration targets (SCCTs) onboardthe rover. For each major element (SiO2, TiO2, Al2O3, FeOT, MgO, CaO, Na2O, K2O), the database was subdivided into five“folds” with similar distributions of the element of interest. One fold was held out as an independent test set, and the remaining fourfolds were used to optimize multivariate regression models relating the spectrum to the composition. We considered a variety of models, and selected several for further investigation for each element, based primarily on the root mean squared error of prediction (RMSEP) on the test set, when analyzed at 3m. In cases with several models of comparable performance at 3 m, we incorporated the SCCT performance at different distances to choose the preferred model. Shortly after landing on Mars and collecting initial spectra of geologic targets, we selected one model per element. Subsequently, with additional data from geologic targets, some models were revised to ensure results that are more consistent with geochemical constraints. The calibration discussed here is a snapshot of an ongoing effort to deliver the most accurate chemical compositions with SuperCam LIBS.

Mars 2020↗

A Modern Load Relief Guidance Scheme for Space Launch Vehicles

Launch vehicle load relief algorithms are concerned with realizing a reduction of transient bending moments near maximum dynamic pressure. Traditional approaches to load relief typically use inner-loop acceleration feedback to reduce the wind-induced angle of attack. When implemented in the inner loop, load relief bandwidth is necessarily limited by the achievable stability margins, and when acceleration feedback is employed, by the uncertainty associated with structural modes that couple with the body-mounted accelerometer. The structure of inner loop load relief increases the dimensionality of the flight control gain and filter optimization problem. Most importantly, classical load relief laws do not take advantage of high-rate and high-accuracy GPS-aided inertial velocity data that is readily available from modern strap down IMUs. In this paper, a novel load relief guidance scheme is described that uses direct angle-of-attack feedback in a clever mechanization. The steering commands are determined by examining the wind-perturbed dynamics of a launch vehicle with respect to a gravity turn ascent trajectory. An angle of attack estimate is derived from GPS-aided inertial data and pre-launch range wind measurements, and it is shown that a reduction worst-case rigid-body loads can be realized without requiring air data. The algorithm also includes a high-rate navigation data preprocessing scheme that operates directly on the IMU delta-theta and delta-velocity measurements in order to produce a filtered acceleration estimate at the vehicle center of mass. The outer-loop guidance scheme simplifies the design process for the classical inner-loop autopilot. Algorithm performance is demonstrated using Monte Carlo analysis of a representative liquid booster in a production high fidelity launch vehicle simulation.

NESC↗

Fast Assessment of Metal Performance through Dislocation Physics and Machine Learning

The microstructure of metals is key to their mechanical properties. The types, density, composition and morphology of crystal defects all have pronounced impact on the properties. Changes to the microstructure occurring during processing and use can be very striking. The emerging technology additive manufacturing (AM) has the potential to improve performance by allowing optimized designs, but the process and environments can lead to unusual microscale features whose properties must be understood and characterized to enable higher technological readiness levels and application. Experimentally, an extensive evaluation of mechanical properties of 3D printed metals is a challenge, and anomalous effects related to the AM process add complexity. We present a new machine learning (ML) model predicting mechanical response based on dislocation mediated plasticity simulations. A large set of 3D discrete dislocation dynamics simulations with wide ranges of loading conditions is transformed to preprocessed data ready for training with the ML model. The trained model can predict the mechanical response of Mo30W for a given microstructure evolution, providing key information essential for optimization of AM processing.

Jaehyun Cho↗

2022 Spring Internship Exit Presentation

As efforts of the National Aeronautics and Space Administration (NASA) and the Federal Aviation Administration (FAA) continue to digitize the air traffic management (ATM) domain, there is countless times of need for downstream natural language processing (NLP) tasks such as named entity recognition, text summarization, classification, and more. Although there are a plethora of open-sourced pre-trained transformer models in the NLP field such as BERT, RoBERTa, XLNet, and GPT-3, these models are trained on general corpora and perform poorly on domain-specific terminology and phraseology seen in ATM documents such as Notice to Airmen (NOTAMs) and Letters of Agreement (LoA). Our proposed research objective will be to first gather a large corpus of air traffic management related documents, orders, notices, books, technical papers, conference papers, articles, and other miscellaneous sources of text data from the FAA, NASA, and accredited conference and publication societies. After gathering this data, many steps will have to be taken to collate and preprocess the data into a format understandable by our test transformer models. Thirdly, we will set up training pipelines to train the RoBERTa model on its unsupervised training task masked language modelling (MLM) using resources provided by the NASA Advanced Supercomputing (NAS) facilities. Finally, these fine-tuned transformer models will be evaluated on their performance on down-stream NLP tasks as mentioned above, to show whether they will be effective when working with ATM related data or not. Once complete, this model could be made open-sourced on the HuggingFace website, where the rest of the ATM community can access and utilize this tool.

NLP↗