Search NASA⌕ Search

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

MERRA-2 Input Observations: Summary and Assessment

The Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA-2) is an atmospheric reanalysis, spanning 1980 through near-realtime, that uses state-of-the-art processing of observations from the continually evolving global observing system. The effectiveness of any reanalysis is a function not only of the input observations themselves, but also of how the observations are handled in the assimilation procedure. Relevant issues to consider include, but are not limited to, data selection, data preprocessing, quality control, bias correction procedures, and blacklisting. As the assimilation algorithm and earth system models are fundamentally fixed in a reanalysis, it is often a change in the character of the observations, and their feedbacks on the system, that cause changes in the character of the reanalysis. It is therefore important to provide documentation of the observing system so that its discontinuities and transitions can be readily linked to discontinuities seen in the gridded atmospheric fields of the reanalysis. With this in mind, this document provides an exhaustive list of the input observations, the context under which they are assimilated, and an initial assessment of selected core observations fundamental to the reanalysis.

Satellite Observations↗

Exploring and Analyzing Climate Variations Online by Using NASA MERRA-2 Data at GES DISC

NASA Giovanni (Goddard Interactive Online Visualization ANd aNalysis Infrastructure) (http:giovanni.sci.gsfc.nasa.govgiovanni) is a web-based data visualization and analysis system developed by the Goddard Earth Sciences Data and Information Services Center (GES DISC). Current data analysis functions include Lat-Lon map, time series, scatter plot, correlation map, difference, cross-section, vertical profile, and animation etc. The system enables basic statistical analysis and comparisons of multiple variables. This web-based tool facilitates data discovery, exploration and analysis of large amount of global and regional remote sensing and model data sets from a number of NASA data centers. Long term global assimilated atmospheric, land, and ocean data have been integrated into the system that enables quick exploration and analysis of climate data without downloading, preprocessing, and learning data. Example data include climate reanalysis data from NASA Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2) which provides data beginning in 1980 to present; land data from NASA Global Land Data Assimilation System (GLDAS), which assimilates data from 1948 to 2012; as well as ocean biological data from NASA Ocean Biogeochemical Model (NOBM), which provides data from 1998 to 2012. This presentation, using surface air temperature, precipitation, ozone, and aerosol, etc. from MERRA-2, demonstrates climate variation analysis with Giovanni at selected regions.

knowledge base↗

Interactive Multi-Instrument Database of Solar Flares

The fundamental motivation of the project is that the scientific output of solar research can be greatly enhanced by better exploitation of the existing solar/heliosphere space-data products jointly with ground-based observations. Our primary focus is on developing a specific innovative methodology based on recent advances in "big data" intelligent databases applied to the growing amount of high-spatial and multi-wavelength resolution, high-cadence data from NASA's missions and supporting ground-based observatories. Our flare database is not simply a manually searchable time-based catalog of events or list of web links pointing to data. It is a preprocessed metadata repository enabling fast search and automatic identification of all recorded flares sharing a specifiable set of characteristics, features, and parameters. The result is a new and unique database of solar flares and data search and classification tools for the Heliophysics community, enabling multi-instrument/multi-wavelength investigations of flare physics and supporting further development of flare-prediction methodologies.

Heliophysics↗

Climatespark: an In-Memory Distributed Computing Framework for Big Climate Data Analytics

The unprecedented growth of climate data creates new opportunities for climate studies, and yet big climate data pose a grand challenge to climatologists to efficiently manage and analyze big data. The complexity of climate data content and analytical algorithms increases the difficulty of implementing algorithms on high performance computing systems. This paper proposes an in-memory, distributed computing framework, ClimateSpark, to facilitate complex big data analytics and time-consuming computational tasks. Chunking data structure improves parallel I/O efficiency, while a spatiotemporal index is built for the chunks to avoid unnecessary data reading and preprocessing. An integrated, multi-dimensional, array-based data model (ClimateRDD) and ETL operations are developed to address big climate data variety by integrating the processing components of the climate data lifecycle. ClimateSpark utilizes Spark SQL and Apache Zeppelin to develop a web portal to facilitate the interaction among climatologists, climate data, analytic operations and computing resources (e.g., using SQL query and Scala/Python notebook). Experimental results show that ClimateSpark conducts different spatiotemporal data queries/analytics with high efficiency and data locality. ClimateSpark is easily adaptable to other big multiple- dimensional, array-based datasets in various geoscience domains.

Hu, Fei↗

Automated Pneumothorax Diagnosis using Deep Neural Networks

Thoracic ultrasound can provide information leading to rapid diagnosis of pneumothorax with improved accuracy over the standard physical examination and with higher sensitivity than anteroposterior chest radiography. However, the clinical We have Furthermore, remote environments, such as the battlefield or deep-space exploration, may lack expertise for diagnosing developed an automated image interpretation pipeline for the analysis of thoracic ultrasound data and the classification of pneumothorax events to provide decision support in such situations. Our pipeline consists of image preprocessing, data augmentation, and deep learning architectures for medical diagnosis. In this work, we demonstrate that robust, accurate interpretation of chest images and video can be achieved using deep neural networks. A number of novel image processing techniques were employed to achieve this result. Affine transformations were applied for data augmentation. Hyperparameters were optimized for learning rate, dropout regularization, batch size, and epoch iteration by a sequential model-based Bayesian approach. In addition, we utilized pretrained architecturesinterpretation of a patient medical image is highly operator dependent. certain pathologies., applying transfer learning and fine-tuning techniques to fully connected layers. Our pipeline yielded binary classification validation accuracies of 98.3% for M-mode images and 99.8% with B-mode video frames.

US Army collaboration↗

The Sentinel-2 MSI Can Increase the Temporal Resolution of 30m Satellite-Derived LAI Estimates

The successful launch of the European Space Agency (ESA) Sentinel-2A (S2-A) on 23 June 2015 with its MultiSpectral Instrument (MSI) provides an important means to augment Earth-observation capabilities following the legacy of Landsat. After the three-month satellite commissioning campaign, the MSI onboard S-2A is performing very well (ESA, 2015). By 3 December 2015, the sensor data records have achieved provisional maturity status and have been accessed in level-1C Top-Of-Atmosphere (TOA) reflectance by the remote sensing community worldwide. Near-nadir observations by the MSI onboard S-2A and the Operational Land Imager (OLI) onboard Landsat 8 were collected during Simultaneous Nadir Overpasses as well as nearly coincident overpasses. This paper presents a processing chain using harmonized S-2A MSI and Landsat 8 OLI sensors to obtain increased temporal resolution in Leaf Area Index (LAI) estimates using the red-edge band B8A of MSI to replace the NIR band B08. Results demonstrate that LAI estimates from the MSI and OLI are comparable, and, given sufficient preprocessing for atmospheric correction and geometric rectification, can be used interchangeably to improve the frequency with which low LAI canopies can be monitored.

Dungan, Jennifer L.↗

Implementing Geometric Surface Imperfections into Sandwich Composite Cylinder Finite Element Method Models

The buckling responses of certain cylindrical shell structures are extremely sensitive to geometric surface imperfections. The NASA Engineering and Safety Center (NESC) Shell Buckling Knockdown Factor Project (SBKF) is conducting research to develop analysis-based buckling design recommendations. Experiments are used to verify the analysis-based factors, but the sensitivity of the test articles to geometric imperfections requires implementing as-manufactured imperfections into high-fidelity finite element method models. Data collection methods such as structured light scanning are used for all geometric surface data used in this work. Common preprocessing and visualization steps used in SBKF are discussed, and steps on how surface scans are prepared for implementation into a finite element model is described. The Python Tool for Implementing Geometric Imperfections in Reduced Structures (Py_TIGIRS), written specifically for the use with SBKF, is briefly described and uses eight functions to extract, modify, and write geometric imperfections into Abaqus input files. Results of the pre-processing methods and results from Py_TIGIRS are provided and compared for Composite Test Article (CTA) 8.2B. Excellent agreement between the visualized scan data and the FEM-extracted geometry is demonstrated. A brief example of why geometric surface imperfections are significant in nonlinear numerical analyses for thin cylinders in axial compression is provided as motivation to use tools such as Py_TIGIRS. Future development of Py_TIGIRS including expansion to structures of arbitrary geometry is planned.

Sandwich structures↗

Entwine Point Tiles for 3D Visualization and Querying of ICESat-2

Point Cloud data from non-optical sensors present challenges in scientific computing in both volume of data and files, even for cloud services environments. As part of the Multi-Mission Algorithm and Analysis Platform (MAAP), a joint open science platform for global biomass modelling, we’ve developed a cloud optimized workflow for using ATL08 (ICESat-2) data as a point cloud. For MAAP, the ATL08 data product is published as Entwine Point Tiles (EPT), allowing users to visualize and query the full extent of this collection interactively without pre-downloading, or preprocessing. The EPT format is a cloud-optimized point cloud data format which re-organizes points into a cloud friendly spatially indexed data structure. MAAP uses AWS S3 to store these point clouds and serves them over OGC specified APIs, 3DTiles for visualization, and WFS for querying. This workflow allows for interactive 3D visualizations in a web browser, including notebook environments and facilitates on the fly subsetting for interactive data exploration, all of which can be applied to other similar sensors.

Alex Mandel↗

A Quantitative Analysis on the Use of Supervised Machine Learning in Earth Science

Recent review papers (Ball et al., 2017; Reichstein et al., 2019) have investigated the opportunities and challenges in applying supervised machine learning (ML) techniques to Earth science problems. A common challenge is the lack of training (or labeled) data. Supervised ML, and especially deep learning (DL), require large training datasets. While there are large, open access Earth science archives, the data typically require preprocessing in preparation for supervised ML, frequently including manual labeling. Our objective is to understand the landscape of supervised ML in the Earth sciences, including which research communities have most rapidly adopted supervised ML, which algorithms are applied, and what data are used to train these algorithms. We conducted a literature survey of Earth science papers published during the last 10 years in journals from the American Geophysical Union (AGU), American Meteorological Society (AMS), the Institute of Electrical and Electronics Engineers(IEEE), and the Society of Photo-Optical Instrumentation Engineers (SPIE). We identified papers containing the terms ML, DL, or the names of individual supervised ML algorithms. "Earth science" is an additional required search term for IEEE and SPIE. We investigate trends in supervised ML usage during the 10-year study period, and manually analyzed AGU papers from 2018-2019 to enable deep-dive statistics.

Katrina S Virts↗

Analysis Ready Satellite Data

Analyais-Ready Data (ARD) specifications have gained rcent prominence in the field of land-related Earth Observations. The ARD label enables users to recognize data that need a minimum of preprocessing before analysis. However, the regular geolocation requirements make Level 2 data in other disciplines problematic. While Level 3 gridded data can satisfy the geolocation requirement, they often sacrifice spatial resolution and other information, such as extreme values. This talk outlines this dilemma with some potential approaches to it.

Analysis-Ready Data↗

Mine Discrimination Using Multispectral Imagery With Feedforward Neural Networks

Simulated mine detection was performed on a polarimetric hyperspectral imaging dataset collected by using an acousto-optic tunable filter camera. A feedforward artificial neural network was programmed to recognize predefined spectral "templates." The simulation results are provided along with the preprocessing steps and window sizes leading to mine detection without false alarms.

polarimetric↗

Application of ML/AI for Identifying Earth Science Datasets in Research Publications

NASA Data Active Archive Centers, or DAACs, ingest, store and distribute data acquired from satellites, ground systems as well as modelling data. These data are organized by the datasets, each presenting collection of files usually associated with the certain mission, instrument, processing level, parameter(s), algorithm and/or model. The number of datasets offered by a single DAAC to the public varies. GES DISC, for example, currently offers for public use approximately ~1,300 datasets. While each publicly offered dataset comes with supporting documentation, it is challenging for novice and even experienced scientists to navigate among the datasets that offer similar parameters to find the datasets for their particular research application. Supplying dataset documentation with the scientific paper citations that refer to that dataset provides means for the dataset users to educate themselves with the application research that dataset is being used in. Collecting citations of the papers that use the datasets for their research yield valuable insights into application areas of those datasets, information about usage of the dataset groups for specific applications and those application topics. It also gives insights into the “deep metrics” of the dataset usage, as opposed to the common metrics of the dataset usage such as number of users who downloaded the dataset files and volumes of downloaded data. Association of a certain scientific paper with the dataset(s) presents a challenge because most of the paper authors do not properly cite the datasets, datasets usually have cryptic names and Digital Object Identifiers (DOIs) that are used for dataset identification were assigned to the datasets only few years ago. Simple Google or online library search do not provide even meaningful fraction of the results when performed by the dataset name or DOI, however they provide too many results when the search is done by more broader terms such as mission and instrument names. Attempts to create an AI system capable to identify dataset in the scientific papers have already been made using neural networks classifiers on the basis of the dataset mission, instrument and variable name. This method was applied to NASA SEDAC, which has 41 datasets in total. In GES DISC there can be as many as ~100 datasets per mission/instrument with some of the datasets consisting of multiple variables so there is a need for more differentiating parameters for dataset identification in the paper. The approach we are currently investigating is creating AI classifiers that are based on multiple dataset features, or keywords, extracted from the NASA Earthdata Common Dataset Repository (CMR). The features are weighted based on how precisely they can identify a dataset. The classifier uses preprocessed paper text as input and searches for the CMR datasets whose feature sets are the closest to the feature sets contained in the paper. The challenges of dataset identification include variety of ways the paper authors describe the datasets in their papers and incomplete tagging of the CMR dataset description (DIFs).

Irina Gerasimov↗

Design and Analysis of Convolutional Neural Network for RF Signal Modulation Classification for In-Orbit Deployment

To effectively transmit data to and from satellites requires a complex and robust RF communication system. Commonly, several different types of signal modulations may be required to maximize satellite efficiency depending on a variety of unexpected channel impairments. We propose a neural network algorithm capable of learning these RF signal modulations using a supervised learning technique designed for low power, high-efficiency in-orbit deployment. The work presented demonstrates a convolutional neural network (CNN) capable of learning and recognizing a set of modulation schemes commonly used to transmit RF information. We are capable of recognizing the modulation scheme from the I and Q data channels directly, with no preprocessing or data conversion required other than breaking the incoming signal into a set of uniform normalized samples. We perform a network design and size analysis, showing that reasonably high accuracy can be obtained using networks with a relatively low number of trainable parameters. Given that a user of a system such as this may wish to receive a signal using a modulation scheme that the network has not previously learned, we demonstrate that transfer learning can learn new modulation schemes by retraining only the fully connected layers in the CNN. Thus, this type of network would excel in outer space deployment using high-efficiency transfer learning hardware. Modulation recognition can be performed through rapid feedforward computation, and the CNN training process is significantly simplified when learning new modulations is required.

CNN↗

DELTA: An Open-Source Framework to Simplify Deep Learning with Satellite Imagery

DELTA (Deep Earth Learning, Tools, and Analysis) is an open-source framework developed at NASA for deep learning on satellite imagery based on tensorflow. It helps simplify data engineering and preprocessing steps and reduces the need for a lot of the boilerplate code that needs written to make datasets palatable for machine learning. This lets data scientists focus on model development while DELTA handles the grunt work. This presentation will demonstrate DELTA’s functionality and share some examples from an active project using it for flood mapping.

Michael von Pohle↗

Classifying Agnostic Biosignatures using Raman, VNIR, and Elemental Data

How can we use our current wealth of terrestrial data, encompassing biogenic and abiogenic systems, to determine the distinguishing properties of life? SCOBI (Statistical Classification of Biosignature Information) uses machine learning techniques to algorithmically identify combinations of measurements that are “indicative of life”. A set of ~1000 observations, comprising elemental abundance, isotopic fractionation, VNIR reflectance, and (in progress) Raman spectra, have been assembled from existing literature and databases. The observations cover systems classified as “indicative alive” (e.g., cells, vegetation), “indicative non-alive” (e.g., fossils, teeth), “mixed indicative” (e.g., soil, pond water), or “non-indicative” (e.g., rocks, meteorites). VNIR data was preprocessed by linear interpolation from 400-2100 nm and smoothed with a Savitzky-Golay filter. To limit the amount of Earth-biochemistry-specific (non-agnostic) information included, the first five spectral features extracted were number of peaks, number of troughs, mean reflectance, mean peak width, and broadest peak width. To help further emphasize agnostic biosignatures, Earth-specific features such as chlorophylls have been manually flagged so that feature importance with and without them can be compared. Classifiers including k-nearest neighbors (KNN), Gaussian Naïve Bayes (GNB), logistic regression (LR), random forest (RF), and support vector machine (SVM) were implemented, as was a combination voting classifier. Performance metrics included false positive rates, false negative rates, and AUC with 50-50 test/train splits (Monte Carlo simulations). Key takeaways from this stage, prior to the inclusion of Raman spectra, are (1) the overall success rate of 0.933 AUC was most heavily influenced by the elemental abundance data; and (2) VNIR reflectance had the lowest classification performance with 0.52 AUC (58% of objects correctly classified). The next steps are to complete integration of Raman spectral data and to improve the approach to pre-processing and feature extraction for both types of spectral data, such as automated baseline removal, whole spectrum matching, and dimensionality reduction.

Biosignatures↗

DELTA: An Open-Source Framework to Simplify Machine Learning with Satellite Imagery

DELTA (Deep Earth Learning, Tools, and Analysis) is an open-source framework developed at NASA to simplify running and training machine learning (ML) models on satellite imagery. Users new to machine learning can run existing ML models on satellite imagery with minimal setup and configuration. For experienced ML users, DELTA helps simplify data engineering, preprocessing steps, and reduces the need for boilerplate code that needs written to make satellite imagery datasets palatable for machine learning. This lets data scientists focus on model development while DELTA handles the imagery manipulation. This presentation will demonstrate DELTA’s functionality and share some examples from an active project using it for flood mapping using imagery from multiple satellite sources

Michael von Pohle↗

Bhutan Water Resources III: Analyzing Forest Disturbances and Climate Data in Bhutan to Create a Tool for Assisting the Himalayan Environment Rhythm Observation and Evaluation Systems (HEROES) Project

Forest disturbances from bark beetle outbreaks are a major concern in Bhutan, known to cause extensive tree mortality to pine and spruce forests. The NASA DEVELOP team partnered with the Ugyen Wangchuck Institute of Conservation and Environmental Research (UWICER), the Bhutan Foundation, and the Karuna Foundation to assess forest changes for the districts of Bumthang and Haa from 2000 to 2018. The project used preprocessed meteorological data from the Climate Hazards Center Infrared Precipitation with Station (CHIRPS) and Famine Early Warning System Network Land Data Assimilation System (FLDAS), along with Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper plus (ETM+), and Landsat 8 Operational Land Imager (OLI) to assess apparent forest disturbance occurrences and observed climate trends. Shuttle Radar Topography Mission (SRTM) was used to resolve variations in elevation and slope for mountainous regions. Using the Google Earth Engine LandTrendr (LT) code algorithm, along with Landsat data, the team developed an app called Forest Disturbances Detection Toolbox (FDDT) to assess forest changes in Bhutan. The app includes climate variables for the focus districts, along with LT variables, which allows the end users to further examine the cause of disturbances. The team compared geocoordinates for known disturbances with LT disturbance detection products. Although additional work is needed in the future to validate the project end products from the FDDT, the tool will be provided to the project partners to aid forest management efforts in Bhutan.

Tashi Choden↗

Fast Assessment of Metal Performance through Dislocation Physics and Machine Learning

The microstructure of metals is key to their mechanical properties. The types, density, composition and morphology of crystal defects all have pronounced impact on the properties. Changes to the microstructure occurring during processing and use can be very striking. The emerging technology additive manufacturing (AM) has the potential to improve performance by allowing optimized designs, but the process and environments can lead to unusual microscale features whose properties must be understood and characterized to enable higher technological readiness levels and application. Experimentally, an extensive evaluation of mechanical properties of 3D printed metals is a challenge, and anomalous effects related to the AM process add complexity. We present a new machine learning (ML) model predicting mechanical response based on dislocation mediated plasticity simulations. A large set of 3D discrete dislocation dynamics simulations with wide ranges of loading conditions is transformed to preprocessed data ready for training with the ML model. The trained model can predict the mechanical response of Mo30W for a given microstructure evolution, providing key information essential for optimization of AM processing.

Jaehyun Cho↗