Search NASA⌕ Search

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Supervised Machine Learning Approach for Classifying Earth Science Publications

The data collections archived and distributed by the GES DISC NASA data center are widely utilized for various Earth Science studies. As these collections are created, many research works are published regarding these collections' algorithms, their validation, and their applications. As NASA data centers collect these publications for public use, it is helpful to categorize them based on how they relate to their associated datasets. Specifically, whether the publication linked to the GES DISC dataset is using it for applicational research, describing the algorithm used for the dataset creation, validating the dataset, or providing a general overview of the data collection. Currently, this process requires simple manual labeling, and as such, it may be possible to solve via automation. To approach this problem, machine learning classifiers were developed to predict a publication's category. Manually labeled publications were used as the training data for the supervised machine learning algorithms, specifically Random Forest and Multinomial Naïve Bayes. After balancing the dataset and implementing the Multinomial Naïve Bayes algorithm, the classification accuracy achieved was substantially higher than the baseline accuracy, thus significantly improving the efficiency of publication labeling.

Rohan Dayal↗

Carbon Storage Technical Viability Approach (CS TVA): An Integrated Approach for Feasibility and Data Resource Assessment

There is currently a poor understanding and lack of workflow to understand the technical viability of carbon storage spatially. To address this gap, the multi-faceted Carbon Storage Technical Viability Approach (CS TVA) is being developed to incorporate CO2 storage resources, environmental and socio-economic justice (EJ/SJ) factors to enable more comprehensive assessments. The CS TVA includes a (1) matrix framework, (2) an integrated and labeled database, (3) a data availability assessment workflow, and (4) spatial data availability assessment results. This approach leverages spatial and data science analytics to communicate data density, uncertainty, and gaps. The workflow can be applied in whole or in part, based on user needs.

Rodriguez, Neyda Cordero↗

Facilitating information transfer in the EOS era

A simple interactive demonstration program has been written in C to allow a user to input data field descriptions as label format. This program generates a full RECFMT (record format) description and the complete transfer syntax description notation (TSDN) file. It is intended that this program be upgraded to operational quality and be made available to users to simplify the description and TSDN file construction task. The total set of capabilities, from the standard formatted data unit packaging of related files and consistent segment structures, through the type definition techniques and the call server, will constitute a unique tool for the systematic transfer of data. This software on each end may be independent, one end from the other. With it available, local software that will be needed to convert user files to and from the canonical interface will be appreciably simplified.

Billingsley, Frederic C.↗

A Weakly Supervised Machine Learning Procedure for Magnet Quench Diagnostics

Voltage taps remain the standard and reliable diagnostic tool for detecting quenches in superconducting magnets. However, they identify a quench only at the time of voltage rise and do not provide information on earlier physical precursors. In this work, we investigate whether acoustic emission data can reveal precursor activity that occurs before conventional voltage detection using machine learning techniques. We introduce an event selection method and a weakly supervised machine learning procedure to learn data-driven criteria for identifying potential acoustic precursors to quenches. Two Convolutional Neural Network (CNN) architectures are trained: one on acoustic sensor events from our selection procedure and one on the Fast Fourier Transforms (FFTs) of these events. Both networks are trained iteratively using confidence-weighted loss functions to associate certain subsets of training data with a precursor label. We evaluate the performance of these models by examining the time distribution of events classified as potential precursors relative to the quench onset. Results indicate that the proposed approach can possibly distinguish acoustic emission events occurring closer to the quench from earlier acoustic activity during ramping, suggesting the potential for flagging quench precursors in acoustic data.

Khan, Maira [Fermilab] (ORCID:0009000891602387)↗

Computer graphics for management: An abstract of capabilities and applications of the EIS system

The Executive Information Services (EIS) system, developed as a computer-based, time-sharing tool for making and implementing management decisions, and including computer graphics capabilities, was described. The following resources are available through the EIS languages: centralized corporate/gov't data base, customized and working data bases, report writing, general computational capability, specialized routines, modeling/programming capability, and graphics. Nearly all EIS graphs can be created by a single, on-line instruction. A large number of options are available, such as selection of graphic form, line control, shading, placement on the page, multiple images on a page, control of scaling and labeling, plotting of cum data sets, optical grid lines, and stack charts. The following are examples of areas in which the EIS system may be used: research, estimating services, planning, budgeting, and performance measurement, national computer hook-up negotiations.

Solem, B. J.↗

Geometric and radiometric characterization of LANDSAT-D thematic mapper and multispectral scanner data

A geometrically raw image of Washington, D.C. was acquired and radiometrically corrected. The data show little of the detector stripping common in earlier MSS images. The radiometrically corrected data have uniform means and standard deviations for the detectors in each band; however, the data for different detectors utilize a different pattern of DN levels, resulting in ubiquitous stripping of 1 DN amplitude. Band-to-band registration was assessed using color composites and small area correlation techniques. The spectral equivalency of the first four bands of the thematic mapper with the four bands of the MSS is being examined. Geometric analysis of the Washington, D.C. scene have started and a generalized routine for examining the contents of the label files and nonvideo data files was implemented. Several discrepancies from the documentation are described. Night scenes and daytime ocean scenes required for radiometric purposes were identified and the data ordered.

Kieffer, H. H.↗

Identification of sea ice types in spaceborne synthetic aperture radar data

This study presents an approach for identification of sea ice types in spaceborne SAR image data. The unsupervised classification approach involves cluster analysis for segmentation of the image data followed by cluster labeling based on previously defined look-up tables containing the expected backscatter signatures of different ice types measured by a land-based scatterometer. Extensive scatterometer observations and experience accumulated in field campaigns during the last 10 yr were used to construct these look-up tables. The classification approach, its expected performance, the dependence of this performance on radar system performance, and expected ice scattering characteristics are discussed. Results using both aircraft and simulated ERS-1 SAR data are presented and compared to limited field ice property measurements and coincident passive microwave imagery. The importance of an integrated postlaunch program for the validation and improvement of this approach is discussed.

Kwok, Ronald↗

High-Fidelity Dataset Generation for Sensor Anomalies in Power Grids using Hardware-in-the-Loop Testbed

Sensor anomalies in power grids can have significant impacts on the operation of the grid due to the increased reliance of the grid operation on data-driven applications. However, there is a lack of datasets that accurately capture these anomalies as many of the anomalies go undetected using the current bad data detectors. High-fidelity labeled datasets are essential for developing robust applications that can detect and mitigate the impacts of anomalies. In this paper, we propose a hardware-in-the-loop testbed model that can emulate the grid behavior with high-fidelity. This testbed is used to inject anomalies at various levels in the grid architecture and generate labeled datasets. These high-fidelity datasets can be used for development and validation of data-driven applications for detection and mitigation of anomalies in grids and other cyber-physical systems.

Hyder, Burhan↗

Analysis of the Westland Data Set

The "Westland" set of empirical accelerometer helicopter data with seeded and labeled faults is analyzed with the aim of condition monitoring. The autoregressive (AR) coefficients from a simple linear model encapsulate a great deal of information in a relatively few measurements; and it has also been found that augmentation of these by harmonic and other parameters call improve classification significantly. Several techniques have been explored, among these restricted Coulomb energy (RCE) networks, learning vector quantization (LVQ), Gaussian mixture classifiers and decision trees. A problem with these approaches, and in common with many classification paradigms, is that augmentation of the feature dimension can degrade classification ability. Thus, we also introduce the Bayesian data reduction algorithm (BDRA), which imposes a Dirichlet prior oil training data and is thus able to quantify probability of error in all exact manner, such that features may be discarded or coarsened appropriately.

Wen, Fang↗

Development and evaluation of an automatic labeling technique for spring small grains

A labeling technique is described which seeks to associate a sampling entity with a particular crop or crop group based on similarity of growing season and temporal-spectral patterns of development. Human analyst provide contextual information, after which labeling decisions are made automatically. Results of a test of the technique on a large, multi-year data set are reported. Grain labeling accuracies are similar to those achieved by human analysis techniques, while non-grain accuracies are lower. Recommendations for improvments and implications of the test results are discussed.

Crist, E. P.↗

Analysis Ready Satellite Data

Analyais-Ready Data (ARD) specifications have gained rcent prominence in the field of land-related Earth Observations. The ARD label enables users to recognize data that need a minimum of preprocessing before analysis. However, the regular geolocation requirements make Level 2 data in other disciplines problematic. While Level 3 gridded data can satisfy the geolocation requirement, they often sacrifice spatial resolution and other information, such as extreme values. This talk outlines this dilemma with some potential approaches to it.

Analysis-Ready Data↗

User's guide: Programs for processing altimeter data over inland seas

The programs described were developed to process GEODYN-formatted satellite altimeter data, and to apply the processed results to predict geoid undulations and gravity anomalies of inland sea areas. These programs are written in standard FORTRAN 77 and are designed to run on the NSESCC IBM 3081(MVS) computer. Because of the experimental nature of these programs they are tailored to the geographical area analyzed. The attached program listings are customized for processing the altimeter data over the Black Sea. Users interested in the Caspian Sea data are expected to modify each program, although the required modifications are generally minor. Program control parameters are defined in the programs via PARAMETER statements and/or DATA statements. Other auxiliary parameters, such as labels, are hard-wired into the programs. Large data files are read in or written out through different input or output units. The program listings of these programs are accompanied by sample IBM job control language (JCL) images. Familiarity with IBM JCL and the TEMPLATE graphic package is assumed.

Au, A. Y.↗

Frosted Tracks

SAND2025-01893O Frosted Tracks is a software tool to group trajectories according to sequences of their behavior. The goal is to start with a very large number of trajectories and identify groups that exhibit similar behavior patterns. The application combines TICC and Metric DBSCAN clustering algorithms for behavioral segmentation and labeling of air/sea trajectory data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

Satellite remote sensing - An integral tool in acquiring global crop production information

Since NASA's program of research concerning remote sensing was initiated in the 1960s, one of its major objectives has been to advance the state-of-the-art in machine processing of satellite acquired multispectral data. Possibilities have been studied regarding a use of these data to identify type, to monitor condition, and to estimate the ontogenetic stage of cultural vegetation. The present investigation provides a review of the state-of-the-art of the technology used to make remote sensing crop production estimates in foreign regions. Attention is given to Landsat data acquisition, aspects of registration and preprocessing, questions of data transformation, data modeling, proportion estimation, labeling, development stage models, crop condition models, and an outlook regarding future developments.

Hall, F. G.↗

Vestibular efferent neurons project to the flocculus

A bilateral projection from the vestibular efferent neurons, located dorsal to the genu of the facial nerve, to the cerebellar flocculus and ventral paraflocculus was demonstrated. Efferent neurons were double-labeled by the unilateral injections of separate retrograde tracers into the labyrinth and into the floccular and ventral parafloccular lobules. Efferent neurons were found with double retrograde tracer labeling both ipsilateral and contralateral to the sites of injection. No double labeling was found when using a fluorescent tracer with non-fluorescent tracers such as horseradish peroxidase (HRP) or biotinylated dextran amine (BDA), but large percentages of efferent neurons were found to be double labeled when using two fluorescent substances including: fluorogold, microruby dextran amine, or rhodamine labeled latex beads. These data suggest a potential role for vestibular efferent neurons in modulating the dynamics of the vestibulo-ocular reflex (VOR) during normal and adaptive conditions.

Non-NASA Center↗

NASA Tech Briefs, May 2005

Topics covered include: Fastener Starter; Multifunctional Deployment Hinges Rigidified by Ultraviolet; Temperature-Controlled Clamping and Releasing Mechanism; Long-Range Emergency Preemption of Traffic Lights; High-Efficiency Microwave Power Amplifier; Improvements of ModalMax High-Fidelity Piezoelectric Audio Device; Alumina or Semiconductor Ribbon Waveguides at 30 to 1,000 GHz; HEMT Frequency Doubler with Output at 300 GHz; Single-Chip FPGA Azimuth Pre-Filter for SAR; Autonomous Navigation by a Mobile Robot; Software Would Largely Automate Design of Kalman Filter; Predicting Flows of Rarefied Gases; Centralized Planning for Multiple Exploratory Robots; Electronic Router; Piezo-Operated Shutter Mechanism Moves 1.5 cm; Two SMA-Actuated Miniature Mechanisms; Vortobots; Ultrasonic/Sonic Jackhammer; Removing Pathogens Using Nano-Ceramic-Fiber Filters; Satellite-Derived Management Zones; Digital Equivalent Data System for XRF Labeling of Objects; Identifying Objects via Encased X-Ray-Fluorescent Materials - the Bar Code Inside; Vacuum Attachment for XRF Scanner; Simultaneous Conoscopic Holography and Raman Spectroscopy; Adding GaAs Monolayers to InAs Quantum-Dot Lasers on (001) InP; Vibrating Optical Fibers to Make Laser Speckle Disappear; Adaptive Filtering Using Recurrent Neural Networks; and Applying Standard Interfaces to a Process-Control Language.

Source record↗

Image Labeler: Label Earth Science Images for Machine Learning

The application of machine learning for image-based classification of earth science phenomena, such as hurricanes, is relatively new. While extremely useful, the techniques used for image-based phenomena classification require storing and managing an abundant supply of labeled images in order to produce meaningful results. Existing methods for dataset management and labeling include maintaining categorized folders on a local machine, a process that can be cumbersome and not scalable. Image Labeler is a fast and scalable web-based tool that facilitates the rapid development of image-based earth science phenomena datasets, in order to aid deep learning application and automated image classification/detection. Image Labeler is built with modern web technologies to maximize the scalability and availability of the platform. It has a user-friendly interface that allows tagging multiple images relatively quickly. Essentially, Image Labeler improves upon existing techniques by providing researchers with a shareable source of tagged earth science images for all their machine learning needs. Here, we demonstrate Image Labeler’s current image extraction and labeling capabilities including supported data sources, spatiotemporal subsetting capabilities, individual project management and team collaboration for large scale projects.

Acharya, Ashish↗

A Hybrid Approach to Labeling Datasets in Earth Science Publications

NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.

Jacob Atkins↗