Search NASASearch

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Using Ensemble Decisions and Active Selection to Improve Low-Cost Labeling for Multi-View Data

This paper seeks to improve low-cost labeling in terms of training set reliability (the fraction of correctly labeled training items) and test set performance for multi-view learning methods. Co-training is a popular multiview learning method that combines high-confidence example selection with low-cost (self) labeling. However, co-training with certain base learning algorithms significantly reduces training set reliability, causing an associated drop in prediction accuracy. We propose the use of ensemble labeling to improve reliability in such cases. We also discuss and show promising results on combining low-cost ensemble labeling with active (low-confidence) example selection. We unify these example selection and labeling strategies under collaborative learning, a family of techniques for multi-view learning that we are developing for distributed, sensor-network environments.

machine learning

The influence of false color infrared display on training field identification

The overall success of large-scale crop inventories of agricultural regions using Landsat multispectral scanner data is highly dependent upon the labeling of training data by analyst/photointerpreters. The principal analyst tool in labeling training data is a false color infrared composite of Landsat bands 4, 5, and 7. In this paper, this color display is investigated and its influence upon classification errors is partially determined.

Coberly, W. A.

Active Learning with Irrelevant Examples

An improved active learning method has been devised for training data classifiers. One example of a data classifier is the algorithm used by the United States Postal Service since the 1960s to recognize scans of handwritten digits for processing zip codes. Active learning algorithms enable rapid training with minimal investment of time on the part of human experts to provide training examples consisting of correctly classified (labeled) input data. They function by identifying which examples would be most profitable for a human expert to label. The goal is to maximize classifier accuracy while minimizing the number of examples the expert must label. Although there are several well-established methods for active learning, they may not operate well when irrelevant examples are present in the data set. That is, they may select an item for labeling that the expert simply cannot assign to any of the valid classes. In the context of classifying handwritten digits, the irrelevant items may include stray marks, smudges, and mis-scans. Querying the expert about these items results in wasted time or erroneous labels, if the expert is forced to assign the item to one of the valid classes. In contrast, the new algorithm provides a specific mechanism for avoiding querying the irrelevant items. This algorithm has two components: an active learner (which could be a conventional active learning algorithm) and a relevance classifier. The combination of these components yields a method, denoted Relevance Bias, that enables the active learner to avoid querying irrelevant data so as to increase its learning rate and efficiency when irrelevant items are present. The algorithm collects irrelevant data in a set of rejected examples, then trains the relevance classifier to distinguish between labeled (relevant) training examples and the rejected ones. The active learner combines its ranking of the items with the probability that they are relevant to yield a final decision about which item to present to the expert for labeling. Experiments on several data sets have demonstrated that the Relevance Bias approach significantly decreases the number of irrelevant items queried and also accelerates learning speed.

Wagstaff, Kiri

Implementation of Machine Learning Methods for Crater-Based Navigation

Terrain Relative Navigation methods require surface feature detectors to gain information from images used to improve on-board state estimates. This paper presents the development of a crater detection method based on Machine Learning that can extract data from optical images with different crater shapes and sizes, under varying lighting conditions. This work includes an automated capability for generating labeled training data and iterative testing of the neural network-based crater detector. Preliminary results are included to quantify the detector’s accuracy compared to a known crater catalog, given a set of real lunar images from the Lunar Reconnaissance Orbiter.

Sofia G Catalan

Parts Quality Management: Direct Part Marking via Data Matrix Symbols for Mission Assurance

A United States Government Accountability Office (GAO) review of twelve NASA programs found widespread parts quality problems contributing to significant cost overruns, schedule delays, and reduced system reliability. Direct part-marking with Data Matrix symbols could significantly improve the quality of inventory control and parts lifecycle management. This paper examines the feasibility of using 15 marking technologies for use in future NASA programs. A structural analysis is based on marked material type, operational environment (e.g., ground, suborbital, orbital), durability of marks, ease of operation, reliability, and affordability. A cost-benefits analysis considers marking technology (data plates, label printing, direct part marking) and marking types (two-dimensional machine-readable, human-readable). Previous NASA parts marking efforts and historical cost data are accounted for, including in-house vs. outsourced marking. Some marking methods are still under development. While this paper focuses on NASA programs, results may be applicable to a variety of industrial environments.

Moss, Chantrice

A Machine Learning-Based Approach to Time-Series Wave Identification in the Solar Wind

The Wind spacecraft has yielded several decades of high-resolution magnetic field data, a large fraction of which displays small-scale structures. In particular, the solar wind is full of wavelike fluctuations that appear in both the field magnitude and its components. The nature of these fluctuations can be tied to the properties of other structures in the solar wind, such as shocks, that have implications for the time evolution of the solar wind. As such, having a large collection of wave events would facilitate further study of the effects that these fluctuations have on solar wind evolution. Given the large volume of magnetic field data available, machine learning is the most practical approach to classifying the myriad small-scale structures observed. To this end, a subset of Wind data is labeled and used as a training set for a multi-branch 1D convolutional neural network aimed at classifying circularly polarized wave modes. Using this algorithm, a preliminary statistical study of one year of data is performed, yielding about 300,000 wave intervals out of about 5,000,000 solar wind intervals. The wave intervals come about more often in the fast solar wind and at higher temperatures, and the number of waves per day is highly periodic. This machine learning-based approach to wave detection has the potential to be a powerful, inexpensive way to catalog waves throughout decades of spacecraft data.

Samuel Fordin

State Identification for Planetary Rovers: Learning and Recognition

A planetary rover must be able to identify states where it should stop or change its plan. With limited and infrequent communication from ground, the rover must recognize states accurately. However, the sensor data is inherently noisy, so identifying the temporal patterns of data that correspond to interesting or important states becomes a complex problem. In this paper, we present an approach to state identification using second-order Hidden Markov Models. Models are trained automatically on a set of labeled training data; the rover uses those models to identify its state from the observed data. The approach is demonstrated on data from a planetary rover platform.

Aycard, Olivier

User's manual for EZPLOT version 5.5: A FORTRAN program for 2-dimensional graphic display of data

EZPLOT is a computer applications program that converts data resident on a file into a plot displayed on the screen of a graphics terminal. This program generates either time history or x-y plots in response to commands entered interactively from a terminal keyboard. Plot parameters consist of a single independent parameter and from one to eight dependent parameters. Various line patterns, symbol shapes, axis scales, text labels, and data modification techniques are available. This user's manual describes EZPLOT as it is implemented on the Ames Research Center, Dryden Research Facility ELXSI computer using DI-3000 graphics software tools.

Garbinski, Charles

Parts Quality Management: Direct Part Marking of Data Matrix Symbol for Mission Assurance

A United States Government Accountability Office (GAO) review of twelve NASA programs found widespread parts quality problems contributing to significant cost overruns, schedule delays, and reduced system reliability. Direct part marking with Data Matrix symbols could significantly improve the quality of inventory control and parts lifecycle management. This paper examines the feasibility of using direct part marking technologies for use in future NASA programs. A structural analysis is based on marked material type, operational environment (e.g., ground, suborbital, Low Earth Orbit), durability of marks, ease of operation, reliability, and affordability. A cost-benefits analysis considers marking technology (label printing, data plates, and direct part marking) and marking types (two-dimensional machine-readable, human-readable). Previous NASA parts marking efforts and historical cost data are accounted for, including inhouse vs. outsourced marking. Some marking methods are still under development. While this paper focuses on NASA programs, results may be applicable to a variety of industrial environments.

Moss, Chantrice

Earth Science Deep Learning: Applications and Lessons Learned

Deep Learning: A subfield of machine learning; Algorithms inspired by function of the brain; Scales with amount of training data; Powerful tool without the need for feature engineering; Suitable for Earth Science applications. Deep Learning for Earth science at MSFC (Marshall Space Flight Center): Phenomena identification; Hurricane intensity (wind speed) estimation; Severe storm (hailstorm) detection; Transverse bands detection; Entity extraction for knowledge graph creation; Ephemeral water detection.

Labeled Data

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, aviation incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts entities relevant to safety analysts. The custom NER model is built by fine-tuning an existing Bidirectional Encoder Representations from Transformers (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from failure-relevant text. This model performs passably, with a weighted average f1 score of 0.33 across entity types, indicating more labeled training data is needed. Extracted entities are used to form a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported using the SAFECOM system. Similar mishaps are manually clustered and reported as single rows within an FMEA. Foreach cluster, we compute frequency, severity, and overall riskin accordance with FAA standards. This methodology can beapplied as part of a broader safety management system totrack trends in mishaps (e.g., likelihood, severity) and discoverknowledge (e.g., causes, effects) that can be utilized to improvesafety outcomes and system performance.

Machine Learning

Scanning Program

SCAN program uses scanning algorithm to locate tokens in line of input data. Tokens can be command words, numbers, data values, labels. Using SCAN subroutines, user extracts tokens from character strings in languages with simple or complex syntax. SCAN thoroughly tested and implemented in NASA's Descent Design System for Shuttle orbiter. SCAN useful for other programs requiring input scanning. SCAN written in FORTRAN 77.

Mattison, W. C.

Improving Acoustic Models by Watching Television

Obtaining sufficient labelled training data is a persistent difficulty for speech recognition research. Although well transcribed data is expensive to produce, there is a constant stream of challenging speech data and poor transcription broadcast as closed-captioned television. We describe a reliable unsupervised method for identifying accurately transcribed sections of these broadcasts, and show how these segments can be used to train a recognition system. Starting from acoustic models trained on the Wall Street Journal database, a single iteration of our training method reduced the word error rate on an independent broadcast television news test set from 62.2% to 59.5%.

Witbrock, Michael J.

Assessing the Biohazard Potential of Putative Martian Organisms for Exploration Class Human Space Missions

Exploration Class missions to Mars will require precautions against potential contamination by any native microorganisms that may be incidentally pathogenic to humans. While the results of NASA's Viking biology experiments of 1976 have been generally interpreted as inconclusive for surface organisms, the possibility of native surface life has never been ruled out and more recent studies suggest that the case for biological interpretation of the Viking Labeled Release data may now be stronger than it was when the experiments were originally conducted. It is possible that, prior to the first human landing on Mars, robotic craft and sample return missions will provide enough data to know with certainty whether or not future human landing sites harbor extant life forms. However, if native life is confirmed, it will be problematic to determine whether any of its species may present a medical risk to astronauts. Therefore, it will become necessary to assess empirically the risk that the planet contains pathogens based on terrestrial examples of pathogenicity and to take a reasonably cautious approach to bio-hazard protection. A survey of terrestrial pathogens was conducted with special emphasis on those pathogens whose evolution has not depended on the presence of animal hosts. The history of the development and implementation of Apollo anticontamination protocol and recent recommendations of the NRC Space Studies Board regarding Mars were reviewed. Organisms can emerge in nature in the absence of indigenous animal hosts and both infectious and non-infectious human pathogens are theoretically possible on Mars. The prospect of Martian surface life, together with the existence of a diversity of routes by which pathogenicity has emerged on Earth, suggests that the possibility of human pathogens on Mars, while low, is not zero. Since the discovery and study of Martian life can have long-term benefits for humanity, the risk that Martian life might include pathogens should not be an obstacle to human exploration. As a precaution, however, it is recommended that EVA suits be decontaminated when astronauts enter surface habitats when returning from field activity and that biosafety protocol approximating laboratory BSL 2 be developed for astronauts working in laboratories on the Martian surface. Quarantine of astronauts and Martian materials arriving on Earth should also be part of a human Mars mission and this and the surface biosafety program should be integral to human expeditions from the earliest stages of the mission planning.

Warmflash, David

The MODIS Aerosol Algorithm: Critical Evaluation and Plans for Collection 6

For ten years the MODIS aerosol algorithm has been applied to measured MODIS radiances to produce a continuous set of aerosol products, over land and ocean. The MODIS aerosol products are widely used by the scientific and applied science communities for variety of purposes that span operational air quality forecasting in estimates o[ clear-sky direct radiative effects over ocean and aerosol-cloud interactions. The products undergo continual evaluation, including self-consistency checks and comparisons with highly accurate ground-based instruments. The result of these evaluation exercises is a quantitative understanding of the strengths and weaknesses of the retrieval, where and when the products are accurate and the situations where and when accuracy degrades. We intend 10 present results of the most recent critical evaluations including the first comparison of the over ocean products against the shipboard aerosol optical depth measurements of the Marine Aerosol Network (MAN), the demonstration of the lack of sensitivity to size parameter in the over land products and identification of residual problems and regional issues. While the current data set is undergoing evaluation, we are preparing for the next data processing, labeled Collection 6. Collection 6 will include transparent Quality Flags, a 3 km aerosol product and the 500m resolution cloud mask used within the aerosol n:bicvu|. These new products and adjustments to algorithm assumptions should provide users with more options and greater control, as they adapt the product for their own purposes.

Remer, Lorraine