Search NASASearch

Engineering topics

Rahul Ramachandran

Publications and source records attributed to Rahul Ramachandran.

At least 37 records · Page 2

Complications of Metadata Curation for NASA Airborne and Field Campaigns, Platforms, and Instruments

The Airborne Data Management Group (ADMG) curates metadata that describe NASA's airborne and field campaigns, platforms and instruments. This activity is vital to building a useful inventory of sub-orbital Earth science data that improves data discovery and access. During the curation process, many metadata issues were identified that required improvement to campaign and data product metadata. In some cases, locating the needed metadata to add to the inventory was a simple process. For other cases, the information was hard to find. In addition, identifying accurate investigation instrument details to add to the inventory was especially complicated because of the variety of definitions used in the Earth science community for the same concepts. One example of this is the concept of instruments' spatial and temporal resolution. The spatial resolution is one of the more difficult elements to curate given the variations in meaning across various disciplines. Clarified definitions are needed to enable consistency of information across campaigns and instruments. In this presentation, we introduce results from a survey of scientists from various fields in which we asked for definitions of spatial and temporal resolution. Our survey results highlight the importance of creating more universally acceptable definitions for certain metadata elements. By curating sub-orbital field campaign and instrument metadata, ADMG is enabling more efficient discovery and access to NASA observations by allowing science data users to search for certain clearly defined criteria and metadata values.

Ashlyn Shirey

Automated Metadata Scoring Approaches for Earth Observation Data

The Common Metadata Repository (CMR) contains metadata records describing NASA’s Earth observation data products which are archived across 12 data centers also known as Distributed Active Archive Centers (DAACs). To ensure that NASA’s data is discoverable, accessible, and usable, the Analysis and Review of CMR (ARC) Team, located at Marshall Space Flight Center, assesses the quality of these metadata records. The ARC team currently uses a combination of automated and manual methods to check metadata records for quality dimensions such as completeness, correctness, and consistency. In addition to these quality assessments, the team is currently exploring various metadata scoring methods in order to provide normalized results across the twelve DAACs. This method is conducted by using automated methods to assess metadata fields and then provide a numeric score, or grade, based on the analysis. To implement this process, two different approaches have been theorized and are currently being explored by the ARC team. This presentation will describe ARC's two proposed methodologies in more detail, and the pros and cons to using these metadata scoring methods.

Jenny Wood

ImageLabler: Labeling and Managing Image Data for Machine Learning in the Earth Sciences

While machine learning techniques for image classification have been around for a long time, storing and managing the vast number of images required as training data is still a problem for scientists. This is especially true for the field of Earth science, where only recently have experts begun using machine learning techniques for image-based phenomena classification. Image Labeler, a fast and scalable cloud-based tagging platform for Earth science images, seeks to improve upon existing methods of managing images and associated metadata, such as maintaining categorized folders of images on a local machine, a process that can be cumbersome and difficult to scale. The platform facilitates rapid development of image-based Earth science phenomena training datasets by allowing scientists to upload their existing imagery as well as extract new samples from open satellite imagery services made available through NASA’s Global Imagery Browse Service (GIBS). Image Labeler also supports GeoTIFF data, with capabilities such as displaying GeoTIFFs on an interactive map, drawing shapefiles over them, and tagging them with additional metadata. This allows scientists to perform spatiotemporal subsetting with geographic information and develop training data more quickly. Built using modern web technologies, Image Labeler includes additional capabilities such as team collaboration for large-scale image tagging projects. Users can download their data in a machine-learning-ready format, allowing scientists to spend time on experimentation rather than on the collection of training data. In this presentation, we demonstrate how Image Labeler seeks to become a one-stop image data management solution for machine learning applications in Earth science.

Ashish Acharya

Phenomena Portal: Large- Scale Visual Exploration of Atmospheric Phenomena

The Earth science community is experiencing a high influx of remote sensing data due to recent advancements in sensor technology. This enables the community to extend their research on a larger scale than ever before. Unfortunately, traditional data processing techniques do not scale well to these new, high volume data sources. State-of-the-art machine learning (ML) pipelines have been proven to overcome these burdens in various other fields but are underexploited within the physical sciences community. Moreover, ML is reliant on labeled data, which is currently sparsely available, owing to the fact that ML adoption is still in the early stages within the Earth and atmospheric science communities. To address these issues, we developed the Phenomena Portal, a visual exploration tool that uses ML to detect various atmospheric phenomena on a global scale. This allows the Earth and atmospheric science communities to view trends of occurrences of phenomena, identify potential relationships between them, and analyze spatiotemporal patterns over time. These detections can also serve as initial labeled data for ML research pertaining to the respective phenomena. The tool also incorporates feedback from subject matter experts to further improve the model detection accuracy, thereby facilitating human-in-the-loop. This presentation will provide an overview of the ML model development and cloud deployment. We also discuss the capabilities of the user interface for displaying the detections.

Muthukumaran Ramasubramanian

ES2Vec: Earth Science Metadata Suggestions and Analogical Reasoning

As the volume of text-based Earth science research grows, it is increasingly possible to discover latent relationships in the literature. However, traditional methodologies are restricted by limited computational capabilities and intractable problem spaces. Advancements in natural language processing (NLP) have allowed us to use an extensive Earth science corpus to create a domain-specific word vector model, Es2Vec, which we have used to surface latent relationships between Earth science concepts and generate improved keyword tags. Earth science metadata keyword assignment is a challenging problem. Dataset curators select appropriate keywords from the Global Change Master Directory (GCMD) set of keywords. The keywords an are integral part of the search and discovery of these datasets. Hence, the selection of keywords is crucial to increasing the discoverability of datasets. Utilizing machine learning techniques, we provide users with automated keyword suggestions to complement manual selection. We trained a machine learning model that leverages the semantic embedding ability of Word2Vec models to process abstracts and suggest relevant keywords. A user interface tool we built to assist data curators in the assignment of such keywords is also described.

word vectors

Public-Private Partnerships to Enable Discovery, Access, and Use of NASA’s Open Earth Science Datasets

Knowledge transfer between public research institutes and private entities is an essential component of the open science movement. While both public and private institutions are making research advances in technologies, organizational boundaries can hinder knowledge transfer. Productive public-private collaboration frameworks are needed to advance research further.

Elizabeth Fancher

A Quantitative Analysis on the Use of Supervised Machine Learning in Earth Science

Recent review papers (Ball et al., 2017; Reichstein et al., 2019) have investigated the opportunities and challenges in applying supervised machine learning (ML) techniques to Earth science problems. A common challenge is the lack of training (or labeled) data. Supervised ML, and especially deep learning (DL), require large training datasets. While there are large, open access Earth science archives, the data typically require preprocessing in preparation for supervised ML, frequently including manual labeling. Our objective is to understand the landscape of supervised ML in the Earth sciences, including which research communities have most rapidly adopted supervised ML, which algorithms are applied, and what data are used to train these algorithms. We conducted a literature survey of Earth science papers published during the last 10 years in journals from the American Geophysical Union (AGU), American Meteorological Society (AMS), the Institute of Electrical and Electronics Engineers(IEEE), and the Society of Photo-Optical Instrumentation Engineers (SPIE). We identified papers containing the terms ML, DL, or the names of individual supervised ML algorithms. "Earth science" is an additional required search term for IEEE and SPIE. We investigate trends in supervised ML usage during the 10-year study period, and manually analyzed AGU papers from 2018-2019 to enable deep-dive statistics.

Katrina S Virts

Pipeline for Applications-Based Data Discovery

From disaster response and mitigation to monitoring water quality or protecting wildlife habitat, satellite Earth observation data can be applied in countless ways to meet pressing needs and benefit society. The crucial first step toward successful data application is data discovery. Potential users often know exactly what data they need--what Earth feature or phenomenon they need to observe, how frequently, and at what resolution or level of accuracy--but may still struggle to discover the existing observations that meet their needs. We have developed a pipeline to connect applications-based users to specific satellites and data collections within NASA's Earth observation program of record that are highly relevant to their data needs. This pipeline combines available information on satellite and instrument measurement characteristics with an innovative machine learning-based approach that identifies instruments that are most relevant to the feature or phenomenon of interest.

Katrina S Virts

Application of Artificial Intelligence for Surface PM2.5 Estimations from Geostationary Satellite and Atmospheric Numerical Model Data

PM2.5, particulate matter (PM) with a diameter less than or equal to 2.5 μm, is emitted from anthropogenic fuel combustion and forest fires. Due to their small size, PM2.5 can penetrate into respiratory systems and cause or exacerbate serious illness. The US Environmental Protection Agency (EPA) regulates the levels of surface PM2.5 but surface monitoring has spatial and temporal limitations. The Aerosol Optical Depth (AOD) retrievals from the Geostationary Operational Environmental Satellite (GOES) missions and meteorological factors can be utilized as an alternative technique to estimate surface PM2.5 levels at a higher spatial and temporal resolution compared to surface monitors. Traditional estimation approaches rely on linear regression techniques and have limitations modeling the nonlinear relationship between the meteorological factors, AOD retrievals, and surface PM2.5. We compare different machine learning techniques and identify the best-suited model that can represent the nonlinearity between the factors affecting PM2.5 levels

Manisha Khatri