Search NASA⌕ Search

SEARCH · Search NASA

Results for “Labeled Datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Hybrid Approach to Labeling Datasets in Earth Science Publications

NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.

Jacob Atkins↗

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda↗

Foundation AI Models for Science

Foundation Models (FM) are AI models that are designed to replace a task or an application specific model. These FM can be applied to many different downstream applications. These FM are trained using self supervised techniques and can be built on any type of sequence data. The use of self supervised learning removes the hurdle for developing a large labeled dataset for training. Most FM use transformer architecture utilizes the notion of self attention which allows the network to model the influence of distant data points to each other both in space and time. The FM models exhibit emergent properties that are induced from the data. FM can be an important tool for science. The scale of these models results in better performance for different downstream applications and these applications show better accuracy over models built from scratch. FM drastically reduces the cost of entry to build different downstream applications both in time and effort. FM for selected science datasets such as optical satellite data, can accelerate applications ranging from data quality monitoring, feature detection and prediction. FM can make it easier to infuse AI into scientific research by removing the training data bottleneck and increasing the use of science data.

Manil Maskey↗

Revolutionizing Earth Science with Generalized AI Models

Foundation Models (FM) are generalized Artificial Intelligence (AI) models that are designed to replace a task or an application-specific model and can be used for many downstream applications. These FM can be built on any sequence data and are trained utilizing self-supervised approaches. The obstacle of creating a sizable labeled dataset for training is removed by using self-supervised learning. Most FM employ transformer design that takes advantage of the idea of self-attention, allowing the network to represent the impact of distant data points on one another in space and time. The FM models show emergent qualities that are induced from the data. FM can become a valuable tool for Earth science researchers. Due to the size of these models, downstream applications built fine-tuning these FM perform better and exhibit greater accuracy than models created from scratch. FM significantly lowers the entry barrier in terms of both the time and effort required to develop various downstream applications. For some scientific datasets, such as optical remote sensing data, FM can speed up processes like classification, object detection and prediction. By eliminating the training data bottleneck and maximizing the usage of science data, FM can make it simpler to integrate AI into scientific research. Initial results for three different FMs will be presented.

Rahul Ramachandran↗

Anomaly Detection for the Roman Space Telescope Wide Field Instrument’s Science Data Processing Pipeline

The Roman Space Telescope (RST) Wide Field Instrument (WFI) will be utilizing a preliminary Science Data Processing (SDP) pipeline during its Integration and Test, and to some extent during Operations, to track basic statistics and identify known features such as cosmic rays, snowballs as well as possible anomalies in raw detector data. In our detectors, these anomalies appear as jumps in the ramp of a readout and are classified as cosmic rays if they appear as a streak or snowballs if they’re more circular. The WFI employs an array of 18 H4RG-10 detectors that collect image samples. Each set of raw frames within a non-destructive exposure is packaged by the SDP pipeline into image cubes for each detector. Each cube is a time series of 4096 × 4096 accumulating pixel frames. The preliminary analysis pipeline is used to locate anomalies in these time-series accumulation frames and identify the type of anomaly, either natural phenomena or detector characteristic. To compare different methods, we’ve implemented both heuristic-based and data-driven methods to identify anomalies. For the heuristic-based approach, we identify snowballs and cosmic rays by the size and shape of outlier pixel clusters between consecutive frames. For data driven methods, we evaluated a Convolutional Neural Network (CNN) model, and more traditional methods like Principal Component Analysis (PCA). CNN is a supervised learning/classification method. Thus, we used a labeled dataset of anomalies to perform segmentation of the image and identify anomalies. We used previously identified cosmic rays and snowballs to measure the accuracy and efficiency of the mentioned approaches. In evaluating these methods, we aim to pick the best fit for the SDP pipeline’s anomaly detection in terms of both performance and runtime.

Paul Horton↗

Image Labeler: A Web Interface to Catalog Earth Science Events

Advances in machine learning (ML) have made it possible to automatically detect Earth science phenomena from satellite imagery. While useful, ML algorithms typically require an extensive dataset containing labeled images for training. Systematic labeling and management of such datasets is quite cumbersome. With this in mind, we present the Image Labeler. Image Labeler is a fast and scalable cloud-based tool that facilitates the rapid development of Earth science event databases, in order to aid automated ML-based image classification.

Case Study↗

Augmented Reality Data Generation for Training Deep Learning Neural Network

One of the major challenges in deep learning is retrieving sufficiently large labeled training datasets, which can become expensive and time consuming to collect. A unique approach to training segmentation is to use Deep Neural Network (DNN) models with a minimal amount of initial labeled training samples. The procedure involves creating synthetic data and using image registration to calculate affine transformations to apply to the synthetic data. The method takes a small dataset and generates a highquality augmented reality synthetic dataset with strong variance while maintaining consistency with real cases. Results illustrate segmentation improvements in various target features and increased average target confidence.

Torres, Gil↗

AssistTaxi: A Comprehensive Dataset for Taxiway Analysis and Autonomous Operations

The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems. This poster presents AssistTaxi, which is a comprehensive novel dataset which is a collection of images for runway and taxiway analysis. The dataset comprises of more than 300,000 frames of diverse and carefully collected data, gathered from Melbourne (MLB) and Grant-Valkaria (X59) general aviation airports. The importance of AssistTaxi lies in its potential to advance autonomous operations, enabling researchers and developers to train and evaluate algorithms for efficient and safe taxiing. Researchers can utilize AssistTaxi to benchmark their algorithms, assess performance, and explore novel approaches for runway and taxiway analysis. Additionally, the dataset serves as a valuable resource for validating and enhancing existing algorithms as well as facilitating innovation in autonomous operations for aviation. We also propose an initial approach to label the dataset using a contour based detection and line extraction technique.

Data Collection↗

AI Foundation Models for Science: An Open Collaborative Initiative

Foundation Models (FMs), AI models designed to replace task-specific models, are increasingly being recognized for their versatility across numerous downstream applications. These models, trained using self-supervised techniques on any type of sequence data, circumvent the need for large annotated datasets, a major bottleneck in traditional AI model development. FMs can be applied to downstream tasks using few-shot learning and fine-tuning, significantly reducing the need for large labeled training datasets and computational resources. However, the development of FMs requires substantial resources, including access to data and compute power, expertise in the latest models, and specialized scientific knowledge for systematic evaluation. It is challenging for a single group to possess all these capabilities. To address this, NASA IMPACT has initiated an open collaborative effort, leveraging partnerships with the private sector and other groups within and outside NASA, to jointly build FMs. The overarching goal is to develop a consistent and collaborative approach to building FMs for high-value science datasets. This initiative has fostered collaboration within NASA and with external partners, including IBM Research, Clark University, DOE’s ORNL, ESA, and USGS. The effort focuses on identifying key datasets with a wide range of downstream applications, pretraining and building FMs using modified transformer architectures, evaluating compute infrastructure needs, and sharing models, pretraining and fine-tuning code, and data with the community. Furthermore, it aims to train the Earth science community to fine-tune these models for various downstream applications. Our initial effort resulted in the creation of a 100 million parameter HLS Geospatial Model within six months, which was released on HuggingFace. We are now expanding our scope to include data from weather and climate models and investigating multimodal models. We invite those interested in participating in this effort to join us by sharing their use cases, expertise, or data.

Rahul Ramachandran↗

Using Machine Learning to Infer Material Properties of Debris Fragments from X-ray Images in the DebriSat Project

The DebriSat project is a collaboration effort with the NASA Orbital Debris Program Office, the U.S. Space Force Space Systems Command Center, The Aerospace Corporation, and the University of Florida. To date, over 200,000 fragments from this ground-based, hypervelocity impact experiment have been collected, and processing is underway to determine their physical characteristics, such as material, shape, color, characteristic length, and average cross-sectional area. The x-ray process is primarily used to identify the location of the fragments and estimated size for extraction, so that these physical characteristics can be assessed. This paper proposes a machine learning-based approach to characterize materials from x-ray images of debris fragments embedded in soft-catch foam used in the DebriSat project. The novel methodology discussed in this paper will highlight the use of x-ray imagery data to characterize these fragments without extraction or a human-in-the-loop. Both supervised and unsupervised machine learning techniques are utilized with this approach to infer the physical parameters of the fragments embedded in the soft-catch foam panels used in the impact experiment based on x-ray images of the foam panels. Additionally, 3D reconstructions of the extracted fragments are created with images taken from two different angles using the structure from motion (SfM) method. The characteristic lengths and shape from the 3D reconstruction, alongside the physical characteristics of the debris, are used in the inference of the material type. To develop and test the approach, a dataset of x-ray images of debris fragments of varying sizes and materials is collected. Supervised learning methods such as convolutional neural networks (CNNs), support vector machines (SVM), decision trees, and random forest classifiers are used due to the high-dimensional feature spaces of the debris and nonlinear decision boundaries for material categorization. Given the limited pre-labeled data of embedded debris materials smaller than 10 mm, unsupervised machine learning techniques such as clustering algorithms and autoencoders are used, in addition to supervised learning methods. The clustering algorithms group similar fragments together based on their physical properties, and autoencoders reduce the dimensionality of the x ray images and extract relevant features. The performance of the proposed approach's is analyzed using a range of statistical methods, including confusion matrices, receiver operating characteristic curves, and precision-recall curves. The results are compared with those obtained using a baseline approach that relies on manual identification and classification of debris fragments. To evaluate the effectiveness of different machine learning methods, statistical tests such as t-tests, ANOVA, and cross-validation are performed, comparing the performance of CNNs, SVMs, clustering algorithms, and autoencoders. Additional analysis needs to be conducted to identify any sources of bias or variability that may affect the results, such as variations in imaging conditions or fragmentation patterns. Other topics explored are limitations, refinements, and the potential use of semi-supervised learning techniques, such as self-training to label unlabeled datasets and co-training using x-ray images taken from two different angles as two different models.

Saik Anam Siam↗

Using Machine Learning to Infer Material Properties of Debris Fragments from X-ray Images in the DebriSat Project

The DebriSat project is a collaboration effort with the NASA Orbital Debris Program Office, the U.S. Space Force Space Systems Command Center, The Aerospace Corporation, and the University of Florida. To date, over 200,000 fragments from this ground-based, hypervelocity impact experiment have been collected, and processing is underway to determine their physical characteristics, such as material, shape, color, characteristic length, and average cross-sectional area. The x-ray process is primarily used to identify the location of the fragments and estimated size for extraction, so that these physical characteristics can be assessed. This paper proposes a machine learning-based approach to characterize materials from x-ray images of debris fragments embedded in soft-catch foam used in the DebriSat project. The novel methodology discussed in this paper will highlight the use of x-ray imagery data to characterize these fragments without extraction or a human-in-the-loop. Both supervised and unsupervised machine learning techniques are utilized with this approach to infer the physical parameters of the fragments embedded in the soft-catch foam panels used in the impact experiment based on x-ray images of the foam panels. Additionally, 3D reconstructions of the extracted fragments are created with images taken from two different angles using the structure from motion (SfM) method. The characteristic lengths and shape from the 3D reconstruction, alongside the physical characteristics of the debris, are used in the inference of the material type. To develop and test the approach, a dataset of x-ray images of debris fragments of varying sizes and materials is collected. Supervised learning methods such as convolutional neural networks (CNNs), support vector machines (SVM), decision trees, and random forest classifiers are used due to the high-dimensional feature spaces of the debris and nonlinear decision boundaries for material categorization. Given the limited pre-labeled data of embedded debris materials smaller than 10 mm, unsupervised machine learning techniques such as clustering algorithms and autoencoders are used, in addition to supervised learning methods. The clustering algorithms group similar fragments together based on their physical properties, and autoencoders reduce the dimensionality of the x ray images and extract relevant features. The performance of the proposed approach's is analyzed using a range of statistical methods, including confusion matrices, receiver operating characteristic curves, and precision-recall curves. The results are compared with those obtained using a baseline approach that relies on manual identification and classification of debris fragments. To evaluate the effectiveness of different machine learning methods, statistical tests such as t-tests, ANOVA, and cross-validation are performed, comparing the performance of CNNs, SVMs, clustering algorithms, and autoencoders. Additional analysis needs to be conducted to identify any sources of bias or variability that may affect the results, such as variations in imaging conditions or fragmentation patterns. Other topics explored are limitations, refinements, and the potential use of semi-supervised learning techniques, such as self-training to label unlabeled datasets and co-training using x-ray images taken from two different angles as two different models.

Saik Anam Siam↗

State Predictor of Classification Cognitive Engine Applied to Channel Fading

This study presents the application of machine learning (ML) to a space-to-ground communication link, showing how ML can be used to detect the presence of detrimental channel fading. Using this channel state information, the communication link can be used more efficiently by reducing the amount of lost data during fading. The motivation for this work is based on channel fading observed during on-orbit operations with NASA's Space Communication and Navigation (SCaN) testbed on the International Space Station (ISS). This paper presents the process to extract a target concept (fading and not-fading) from the raw data. The pre-processing and data exploration effort is explained in detail, with a list of assumptions made for parsing and labelling the dataset. The model selection process is explained, specifically emphasizing the benefits of using an ensemble of algorithms with majority voting for binary classification of the channel state. Experimental results are shown, highlighting how an end-to-end communication system can utilize knowledge of the channel fading status to identity fading and take appropriate action. With a laboratory testbed to emulate channel fading, the overall performance is compared to standard adaptive methods without fading knowledge, such as adaptive coding and modulation.

Fading↗

Deploying a Self-Supervised Learning Based Model to Search Events Across Space and Time

Motivation - Scientific Study of natural events, phenomena, or disasters require examples which span across time and space. - Machine Learning adaptation is on the rise, but there’s a lack of labeled training datasets that could be used to train or validate the models. - Best case scenario: - There’s an event database that tracks events available through time and space. - Provides all data associated with the events. - Real life scenario: - Some events are better tracked than others. - Scientists need to spend significant time identifying and gathering examples of events from different sources.

Iksha Gurung↗

Image Labeler: Label Earth Science Images for Machine Learning

The application of machine learning for image-based classification of earth science phenomena, such as hurricanes, is relatively new. While extremely useful, the techniques used for image-based phenomena classification require storing and managing an abundant supply of labeled images in order to produce meaningful results. Existing methods for dataset management and labeling include maintaining categorized folders on a local machine, a process that can be cumbersome and not scalable. Image Labeler is a fast and scalable web-based tool that facilitates the rapid development of image-based earth science phenomena datasets, in order to aid deep learning application and automated image classification/detection. Image Labeler is built with modern web technologies to maximize the scalability and availability of the platform. It has a user-friendly interface that allows tagging multiple images relatively quickly. Essentially, Image Labeler improves upon existing techniques by providing researchers with a shareable source of tagged earth science images for all their machine learning needs. Here, we demonstrate Image Labeler’s current image extraction and labeling capabilities including supported data sources, spatiotemporal subsetting capabilities, individual project management and team collaboration for large scale projects.

Acharya, Ashish↗

FloodPlanet: High-Resolution Commercial Imagery for Training and Validation of Deep Learning-Based Models of Inundation Extent

Flooding events are becoming increasingly frequent worldwide and are known to cause extensive damage. Public optical and radar satellite imagery can be used to detect large areas of inundation in rural areas, however, long revisit times and coarse spatial resolution limit applications for short-lived events and urban areas. Commercial constellations such as those operated by Planet offer increased spatial and temporal resolution and can supplement mapping efforts to provide more information to disaster response, relief, and mitigation efforts. Deep learning requires high quality labeled data for training across coincident sensors. The FloodPlanet dataset presented here contains labeled surface water for 18 events across the world based on Planetscope imagery with coincident Harmonized Landsat Sentinel-2 ( HLS) or Sentinel-1 and builds upon the previously existing Sen1Floods11, xBD, and NASA Sentinel-1 datasets. Sen1Floods11 includes 4,831 512x512 pixel overlapping tiles of coincident Sentinel-1 and Sentinel-2 data observing 11 flood events across the world from 2017-2019. The dataset contains a combination of automated and hand-labeled surface water for use in training and validation of inundation modeling efforts. The xBD dataset identifies flood-damaged buildings and indicates the scale of damage to each (none, minor, moderate, and major) from four flood events which occurred in the United States, India, Nepal, and Bangladesh from the same time period. The NASA dataset contains hand-labeled water bodies observed in Sentinel-1 imagery during five flood events within the 2017-2019 period. The effort presented here utilizes observations from these previously investigated flood events to generate labels of surface water at the 3-5m spatial resolution provided by Planetscope and facilitate the comparison between public and commercial data. A data pipeline was built which uses clustering algorithms to pick the most suitable overlapping chips between the public data and PlanetScope data for manual labeling. Labels were created manually using NASA’s ImageLabeler tool and include areas of high- and low-confidence water. The high confidence designation is reserved for areas of open, unobstructed water while low confidence is used for areas of suspected water beneath vegetation, clouds, or cloud shadows. Expected to be released in late 2022, the FloodPlanet dataset will include tiled imagery with a unique ID for each 1024x1024 pixel tile, 7 bands of HLS data, and high- and low-confidence flood labels in both shapefile and tiff formats. The authors will follow Spatial Temporal Access Catalog (STAC) guidelines to release FloodPlanet on the Radiant Earth ML hub, which hosts public datasets for machine learning.

Alexander Melancon↗

Expanding NeMO-Net Machine Learning Capabilities for Citizen Science

NASA NeMO-Net, the neural multi-modal observation and training network for global coral reef assessment, is an open-source deep convolutional neural network and interactive active learning training software aiming to accurately assess the present and past dynamics of coral reef ecosystems through determination of percent living cover and morphology as well as mapping of spatial distribution. We present an interactive citizen science video game, released this April, for desktop and iOS devices where users interactively label morphology classifications over mm-scale 3D coral reef imagery captured using diver photomosaic imagery, the UAV enabled NASA FluidCam instrument, and satellite datasets. To date, the application has had over 40,000 downloads and over 60,000 unique coral reef classifications, each filtered through a user-based rating and expert evaluation system. We also present results from NeMO-Net’s convolutional neural network (CNN) models used to semantically segment 2D satellite imagery as well as projections of 3D coral reconstructions using user input data as training datasets. Fusing datasets using machine learning from multiple remote sensing platforms presents novel methodologies for assessing the health of coral ecosystems, which are critically endangered by a changing climate. In partnering with Mission Blue, the National Oceanic and Atmospheric Administration (NOAA), and the Living Oceans Foundation (LOF), NeMO-Net leverages an international consortium of subject matter experts to provide both proper training for citizen scientists and the generation of a labeled datasets to ingest into machine learning algorithms for global coral reef identification.

NeMO-Net↗