Search NASA⌕ Search

SEARCH · Search NASA

Results for “label quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Learning instrument invariant characteristics for generating high-resolution global coral reef maps

Coral reefs are one of the most biologically complex and diverse ecosystems within the shallow marine environment. Unfortunately, these underwater ecosystems are threatened by a number of anthropogenic challenges, including ocean acidification and warming, overfishing, and the continued increase of marine debris in oceans. This requires a comprehensive assessment of the world's coastal environments, including a quantitative analysis on the health and extent of coral reefs and other associated marine species, as a vital Earth Science measurement. However, limitations in observational and technological capabilities inhibit global sustained imaging of the marine environment. Harmonizing multimodal data sets acquired using different remote sensing instruments presents additional challenges, thereby limiting the availability of good quality labeled data for analysis. In this work, we develop a deep learning model for extracting domain invariant features from multimodal remote sensing imagery and creating high-resolution global maps of coral reefs by combining various sources of imagery and limited hand-labeled data available for certain regions. This framework allows us to generate, for the first time, coral reef segmentation maps at 2-meter resolution, which is a significant improvement over the kilometer-scale state-of-the-art maps. Additionally, this framework doubles accuracy and IoU metrics over baselines that do not account for domain invariance.

Domain Adaptation↗

FloodPlanet: High-Resolution Commercial Imagery for Training and Validation of Deep Learning-Based Models of Inundation Extent

Flooding events are becoming increasingly frequent worldwide and are known to cause extensive damage. Public optical and radar satellite imagery can be used to detect large areas of inundation in rural areas, however, long revisit times and coarse spatial resolution limit applications for short-lived events and urban areas. Commercial constellations such as those operated by Planet offer increased spatial and temporal resolution and can supplement mapping efforts to provide more information to disaster response, relief, and mitigation efforts. Deep learning requires high quality labeled data for training across coincident sensors. The FloodPlanet dataset presented here contains labeled surface water for 18 events across the world based on Planetscope imagery with coincident Harmonized Landsat Sentinel-2 ( HLS) or Sentinel-1 and builds upon the previously existing Sen1Floods11, xBD, and NASA Sentinel-1 datasets. Sen1Floods11 includes 4,831 512x512 pixel overlapping tiles of coincident Sentinel-1 and Sentinel-2 data observing 11 flood events across the world from 2017-2019. The dataset contains a combination of automated and hand-labeled surface water for use in training and validation of inundation modeling efforts. The xBD dataset identifies flood-damaged buildings and indicates the scale of damage to each (none, minor, moderate, and major) from four flood events which occurred in the United States, India, Nepal, and Bangladesh from the same time period. The NASA dataset contains hand-labeled water bodies observed in Sentinel-1 imagery during five flood events within the 2017-2019 period. The effort presented here utilizes observations from these previously investigated flood events to generate labels of surface water at the 3-5m spatial resolution provided by Planetscope and facilitate the comparison between public and commercial data. A data pipeline was built which uses clustering algorithms to pick the most suitable overlapping chips between the public data and PlanetScope data for manual labeling. Labels were created manually using NASA’s ImageLabeler tool and include areas of high- and low-confidence water. The high confidence designation is reserved for areas of open, unobstructed water while low confidence is used for areas of suspected water beneath vegetation, clouds, or cloud shadows. Expected to be released in late 2022, the FloodPlanet dataset will include tiled imagery with a unique ID for each 1024x1024 pixel tile, 7 bands of HLS data, and high- and low-confidence flood labels in both shapefile and tiff formats. The authors will follow Spatial Temporal Access Catalog (STAC) guidelines to release FloodPlanet on the Radiant Earth ML hub, which hosts public datasets for machine learning.

Alexander Melancon↗

Scatter-Reducing Sounding Filtration Using a Genetic Algorithm and Mean Monthly Standard Deviation

Retrieval algorithms like that used by the Orbiting Carbon Observatory (OCO)-2 mission generate massive quantities of data of varying quality and reliability. A computationally efficient, simple method of labeling problematic datapoints or predicting soundings that will fail is required for basic operation, given that only 6% of the retrieved data may be operationally processed. This method automatically obtains a filter designed to reduce scatter based on a small number of input features. Most machine-learning filter construction algorithms attempt to predict error in the CO2 value. By using a surrogate goal of Mean Monthly STDEV, the goal is to reduce the retrieved CO2 scatter rather than solving the harder problem of reducing CO2 error. This lends itself to improved interpretability and performance. This software reduces the scatter of retrieved CO2 values globally based on a minimum number of input features. It can be used as a prefilter to reduce the number of soundings requested, or as a post-filter to label data quality. The use of the MMS (Mean Monthly Standard deviation) provides a much cleaner, clearer filter than the standard ABS(CO2-truth) metrics previously employed by competitor methods. The software's main strength lies in a clearer (i.e., fewer features required) filter that more efficiently reduces scatter in retrieved CO2 rather than focusing on the more complex (and easily removed) bias issues.

Mandrake, Lukas↗

A Machine Learning-Based Cloud Detection and Thermodynamic Phase Classification Algorithm using Passive Spectral Observations

We trained two Random Forest (RF) machine-learning models for cloud mask and cloud thermodynamic phase detection using spectral observations from VIIRS on Suomi NPP (SNPP). Observations from CALIOP were carefully selected to provide reference labels. The two RF models were trained for all-day and daytime-only conditions using a 4-year collocated VIIRS/CALIOP dataset from 2013 to 2016. Due to the orbit difference, the collocated CALIOP and SNPP VIIRS training samples cover a broad viewing zenith angle range, which is a great benefit to overall model performance. The all-day model uses 3 VIIRS infrared (IR) bands (8.6,11, and 12 μm) and the daytime model uses 5 Near-IR (NIR) and Shortwave-IR (SWIR) bands (0.86, 1.24, 1.38, 1.64 and 2.25 μm) together with the 3 IR bands to detect clear, liquid water, and ice cloud pixels. Up to 7 surface types, namely, ocean/water, forest, cropland, grassland, snow/ice, barren/desert, and shrubland, were considered separately to enhance performance for both models. Detection of cloudy pixels and thermodynamic phase with the two RF models were compared against collocated CALIOP products from 2017. It is shown that, with a conservative screening process that excludes the most challenging cloudy pixels for passive remote sensing, the two RF models have high accuracy rates in comparison with the CALIOP reference for both cloud detection and thermodynamic phase. Other existing SNPP VIIRS and Aqua MODIS cloud mask and phase products are also evaluated, with results showing that the two RF models and the MODIS MYD06 optical property phase product are the top 3 algorithms with respect to lidar observations during the daytime. During the nighttime, the RF all-day model works best for both cloud detection and phase, in particular for pixels over snow/ice surfaces. The present RF models can be extended to other similar passive instruments if training samples can be collected from CALIOP or other lidars. However, the quality of reference labels and potential sampling issues that may impact model performance would need further attention.

cloud detection↗

In Situ Water Quality Data for the Chesapeake Bay

This paper examines in situ water quality datameasured during2020-2021in the Chesapeake Bay for comparison with optical satellite data. Thiscollection was performed as part of a NASA project aiming to develop new methods for water quality monitoring from satellite remote sensingusing artificial intelligence. Our objective is to use insitu data as ground-truth to provide water quality classifications, or labels,to their overlapping (in time and location)satellite imagery. Having such labeled data, can help us achieve our project’s longer-termgoal:to train artificial intelligencemodelsto recognize features in spectral informationfor monitoringwater qualityfrom satellites. Because routine monitoring by state agencies is conducted at discrete locations, we obtained a flow-through system operated from small boats to measure waterquality parameters along transects for comparison with two-dimensional maps collected from space, with an initial focus on low oxygenevents, due to their large spatial extent and regular occurrence each summer.We also evaluated similar in situ data collected during 1984-2021by the Chesapeake Program.

Nargess Memarsadeghi↗

AI4MARS: A Dataset for Terrain-Aware Autonomy on Mars

Deep learning has quickly become a necessity for selfdriving vehicles on Earth. In contrast, the self-driving vehicles on Mars, including NASA’s latest rover, Perseverance, which is planned to land on Mars in February 2021, are still driven by classical machine vision systems. Deep learning capabilities, such as semantic segmentation and object recognition, would substantially benefit the safety and productivity of ongoing and future missions to the red planet. To this end, we created the first large-scale dataset, AI4Mars, for training and validating terrain classification models for Mars, consisting of ~326K semantic segmentation full image labels on 35K images from Curiosity, Opportunity, and Spirit rovers, collected through crowdsourcing. Each image was labeled by ~10 people to ensure greater quality and agreement of the crowdsourced labels. It also includes ~1.5K validation labels annotated by the rover planners and scientists from NASA’s MSL (Mars Science Laboratory) mission, which operates the Curiosity rover, and MER (Mars Exploration Rovers) mission, which operated the Spirit and Opportunity rovers. We trained a DeepLabv3 model on the AI4Mars training dataset and achieved over 96% overall classification accuracy on the test set. The dataset is made publicly available.1

Ono, Hiro↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of a macromolecules purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals will show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "bits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Minamitani, Elizabeth Forsythe↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically can not reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "hits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Pusey, Marc L.↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, 51%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear hits. Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Pusey, Marc L.↗

Fluorescent Applications to Crystallization

By covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and tests with model proteins have shown that labeling u to 5 percent of the protein molecules does not affect the X-ray data quality obtained . The presence of the trace fluorescent label gives a number of advantages. Since the label is covalently attached to the protein molecules, it "tracks" the protein s response to the crystallization conditions. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a darker background. Non-protein structures, such as salt crystals, do not show up under fluorescent illumination. Crystals have the highest protein concentration and are readily observed against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. Preliminary tests, using model proteins, indicates that we can use high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that more rapid amorphous precipitation kinetics may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Experiments are now being carried out to test this approach using a wider range, of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.↗

The MODIS Aerosol Algorithm: Critical Evaluation and Plans for Collection 6

For ten years the MODIS aerosol algorithm has been applied to measured MODIS radiances to produce a continuous set of aerosol products, over land and ocean. The MODIS aerosol products are widely used by the scientific and applied science communities for variety of purposes that span operational air quality forecasting in estimates o[ clear-sky direct radiative effects over ocean and aerosol-cloud interactions. The products undergo continual evaluation, including self-consistency checks and comparisons with highly accurate ground-based instruments. The result of these evaluation exercises is a quantitative understanding of the strengths and weaknesses of the retrieval, where and when the products are accurate and the situations where and when accuracy degrades. We intend 10 present results of the most recent critical evaluations including the first comparison of the over ocean products against the shipboard aerosol optical depth measurements of the Marine Aerosol Network (MAN), the demonstration of the lack of sensitivity to size parameter in the over land products and identification of residual problems and regional issues. While the current data set is undergoing evaluation, we are preparing for the next data processing, labeled Collection 6. Collection 6 will include transparent Quality Flags, a 3 km aerosol product and the 500m resolution cloud mask used within the aerosol n:bicvu|. These new products and adjustments to algorithm assumptions should provide users with more options and greater control, as they adapt the product for their own purposes.

Remer, Lorraine↗

Fluorescent Approaches to High Throughput Crystallography

We have shown that by covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and the presence of the probe at low concentrations does not affect the X-ray data quality or the crystallization behavior. The presence of the trace fluorescent label gives a number of advantages when used with high throughput crystallizations. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a dark background. Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Brightly fluorescent crystals are readily found against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. We are now testing the use of high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that kinetics leading to non-structured phases may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Preliminary experiments with test proteins have resulted in the extraction of a number of crystallization conditions from screening outcomes based solely on the presence of bright fluorescent regions. Subsequent experiments will test this approach using a wider range of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.↗

Remote sensing data processing: Two years ago, today, and two years from today

Beginning with a survey of the state-of-the-art of processing remotely sensed data in early 1975, significant developments between that time and the present are chronicled, and technologies for early 1979 are projected. Current technical issues discussed include: training selection and labeling; classification and mensuration; use of satellite indicators to supplement predictions; small scale field structures; physical factors; ancillary data; geometric quality; and the cost of processing.

Holmes, Q. A.↗

Accuracy assessment in the Large Area Crop Inventory Experiment

The Accuracy Assessment System (AAS) of the Large Area Crop Inventory Experiment (LACIE) was responsible for determining the accuracy and reliability of LACIE estimates of wheat production, area, and yield, made at regular intervals throughout the crop season, and for investigating the various LACIE error sources, quantifying these errors, and relating them to their causes. Some results of using the AAS during the three years of LACIE are reviewed. As the program culminated, AAS was able not only to meet the goal of obtaining accurate statistical estimates of sampling and classification accuracy, but also the goal of evaluating component labeling errors. Furthermore, the ground-truth data processing matured from collecting data for one crop (small grains) to collecting, quality-checking, and archiving data for all crops in a LACIE small segment.

Houston, A. G.↗

Adaptable Constrained Genetic Programming: Extensions and Applications

An evolutionary algorithm applies evolution-based principles to problem solving. To solve a problem, the user defines the space of potential solutions, the representation space. Sample solutions are encoded in a chromosome-like structure. The algorithm maintains a population of such samples, which undergo simulated evolution by means of mutation, crossover, and survival of the fittest principles. Genetic Programming (GP) uses tree-like chromosomes, providing very rich representation suitable for many problems of interest. GP has been successfully applied to a number of practical problems such as learning Boolean functions and designing hardware circuits. To apply GP to a problem, the user needs to define the actual representation space, by defining the atomic functions and terminals labeling the actual trees. The sufficiency principle requires that the label set be sufficient to build the desired solution trees. The closure principle allows the labels to mix in any arity-consistent manner. To satisfy both principles, the user is often forced to provide a large label set, with ad hoc interpretations or penalties to deal with undesired local contexts. This unfortunately enlarges the actual representation space, and thus usually slows down the search. In the past few years, three different methodologies have been proposed to allow the user to alleviate the closure principle by providing means to define, and to process, constraints on mixing the labels in the trees. Last summer we proposed a new methodology to further alleviate the problem by discovering local heuristics for building quality solution trees. A pilot system was implemented last summer and tested throughout the year. This summer we have implemented a new revision, and produced a User's Manual so that the pilot system can be made available to other practitioners and researchers. We have also designed, and partly implemented, a larger system capable of dealing with much more powerful heuristics.

Janikow, Cezary Z.↗

An operational ASDAR system

The story of the Aircraft to Satellite Data Relay (ASDAR) program began when airline meteorologists realized that B-747's and other commercial jets provided cockpit displays of digital values for outside air temperature and winds. Later, when a few B-747's were used to carry portable air quality monitoring equipment for the Global Air Sampling Program (GASP), scientists at NASA-Lewis explored ways in which these digital values could be used to label data collected during the GASP flights. Digital values of GASP analyses were recorded along with digital values of location and altitude, time, winds, and temperature, obtained by microprocessors from within the host aircraft's avionics. These data suggested a way in which manually recorded in-flight meteorological reports could be replaced by an automatic system, which could record winds and air temperatures as often as desired. NASA's prototype ASDAR showed that automated data relay by meteorological geostationary satellites could be accomplished from an aircraft. Testing of the instruments and analyses of its data are examined.

Sparkman, James K., Jr.↗

Facilitating information transfer in the EOS era

A simple interactive demonstration program has been written in C to allow a user to input data field descriptions as label format. This program generates a full RECFMT (record format) description and the complete transfer syntax description notation (TSDN) file. It is intended that this program be upgraded to operational quality and be made available to users to simplify the description and TSDN file construction task. The total set of capabilities, from the standard formatted data unit packaging of related files and consistent segment structures, through the type definition techniques and the call server, will constitute a unique tool for the systematic transfer of data. This software on each end may be independent, one end from the other. With it available, local software that will be needed to convert user files to and from the canonical interface will be appreciably simplified.

Billingsley, Frederic C.↗

An Investigation Into HPLC Data Quality Problems

This report summarizes the analyses and results produced by a five-member investigative team of Government, university, and industry experts, established by NASA HQ. The team examined data quality problems associated with high performance liquid chromatography (HPLC) analyses of pigment concentrations in seawater samples produced by the San Diego State University (SDSU) Center for Hydro-Optics and Remote Sensing (CHORS). This report shows CHORS did not validate the methods used before placing them into service to analyze field samples for NASA principal investigators (PIs), even though the HPLC literature contained easily accessible method validation procedures, and the importance of implementing them, more than a decade ago. In addition, there were so many sources of significant variance in the CHORS methodologies, that the HPLC system rarely operated within performance criteria capable of producing the requisite data quality. It is the recommendation of the investigative team to a) not correct the data, b) make all the data that was temporarily sequestered available for scientific use, and c) label the affected data with an appropriate warning, e.g., "These data are not validated and should not be used as the sole basis for a scientific result, conclusion, or hypothesis--independent corroborating evidence is required."

Hooker, Stanford B.↗