Search NASA⌕ Search

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Mars Terrain Segmentation with Less Labels

Planetary rover systems need to perform terrain segmentation to identify drivable areas as well as identify specific types of soil for sample collection. The latest Martian terrain segmentation methods rely on supervised learning which is very data hungry and difficult to train where only a small number of labeled samples are available. Moreover, the semantic classes are defined differently for different applications (e.g., rover traversal vs. geological) and as a result the network has to be trained from scratch each time, which is an inefficient use of resources. This research proposes a semi-supervised learning framework for Mars terrain segmentation where a deep segmentation network trained in an unsupervised manner on unlabeled images is transferred to the task of terrain segmentation trained on few labeled images. The network incorporates a backbone module which is trained using a contrastive loss function and an output atrous convolution module which is trained using a pixel-wise cross-entropy loss function. Evaluation results using the metric of segmentation accuracy show that the proposed method with contrastive pre-training outperforms plain supervised learning by 2%-10%. Moreover, the proposed model is able to achieve a segmentation accuracy of 91.1% using only 161 training images (1% of the original dataset) compared to 81.9% with plain supervised learning.

Wilson, Brian D↗

The Land Surface in Current and Planned MERRA Reanalysis Products

Current global atmospheric reanalysis products such as the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5), the NASA Modern-Era Retrospective analysis for Research and Applications version 2 (MERRA-2), and the Japanese Reanalysis for Three Quarters of a Century (JRA-3Q) provide estimates of land surface states and fluxes, including soil moisture, soil temperature, snow mass, latent and sensible heat fluxes, and runoff, that are widely used in research and applications. These land surface estimates are based on land surface process models and, depending on the reanalysis product, on precipitation observations or the assimilation of land surface observations of soil moisture, soil temperature, snow conditions, and screen-level air temperature and humidity from satellite observations and in situ measurements. In this presentation, we review the land surface modeling and data assimilation components of the suite of current and planned MERRA reanalysis products. In addition to MERRA-2, we will discuss the latest NASA reanalysis, MERRA for the 21st century (M21C), which is currently under production, as well as the development and planning of the next version of the MERRA reanalysis, tentatively labeled MERRA-3. In MERRA-2, observations-based precipitation data products are used to correct the precipitation falling on the land surface. Outside of the high-latitudes and Africa, the daily, 0.5-degree, gauge-based Climate Prediction Center (CPC) Unified (CPCU) product is used. In Africa, the pentad, 2.5-degree, satellite- and gauge-based CPC Merged Analysis of Precipitation (CMAP) product is used. Poleward of 62.5 degrees latitude, the land surface sees the precipitation generated by the atmospheric model in the cycling data assimilation system. This configuration provides improved soil moisture estimates compared to those of the original (version 1) MERRA estimates, which did not benefit from the use of precipitation observations. Moreover, the use of precipitation observations facilitates a seamless spin-up of the land surface initial conditions across the MERRA-2 production streams. The use of a gauge-only precipitation product in MERRA-2 across much of the globe, however, adversely impacts the quality of the MERRA-2 land surface estimates in regions with poor gauge coverage, including most of South America and Australia. Therefore, the forthcoming M21C reanalysis uses satellite- and gauge-based precipitation from the Integrated Multi-satellitE Retrievals for the Global Precipitation Measurement Mission (IMERG). This change results in significant improvements in the quality of the M21C soil moisture estimates in the Southern Hemisphere compared to those from MERRA-2. Planning for MERRA-3 focuses on the assimilation of soil moisture observations from the Soil Moisture Active Passive (SMAP) mission and the Advanced Scatterometer (ASCAT), along with snow cover area fraction observations from the Moderate Resolution Imaging Spectroradiometer (MODIS) to further improve the quality of the land surface estimates from the reanalysis. As a first step towards the assimilation of land surface observations in MERRA-3, the offline (land-only) M21C-Land reanalysis is currently under development as a supplemental M21C product that includes the assimilation of SMAP, ASCAT, and MODIS observations. Preliminary results from M21C and M21C-Land will be discussed in the context of MERRA-2 and plans for MERRA-3.

Rolf Reichle↗

Fluorescent Approaches to High Throughput Crystallography

We have shown that by covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and the presence of the probe at low concentrations does not affect the X-ray data quality or the crystallization behavior. The presence of the trace fluorescent label gives a number of advantages when used with high throughput crystallizations. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a dark background. Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Brightly fluorescent crystals are readily found against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. We are now testing the use of high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that kinetics leading to non-structured phases may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Preliminary experiments with test proteins have resulted in the extraction of a number of crystallization conditions from screening outcomes based solely on the presence of bright fluorescent regions. Subsequent experiments will test this approach using a wider range of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.↗

Water Across Synthetic Aperture Radar Data (WASARD): SAR Water Body Classification for the Open Data Cube

The detection of inland water bodies from Synthetic Aperture Radar (SAR) data provides a great advantage over water detection with optical data, since SAR imaging is not impeded by cloud cover. Traditional methods of detecting water from SAR data involves using thresholding methods that can be labor intensive and imprecise. This paper describes Water Across Synthetic Aperture Radar Data (WASARD): a method of water detection from SAR data which automates and simplifies the thresholding process using machine learning on training data created from Geoscience Australia’s WOFS algorithm. Of the machine learning models tested, the Linear Support Vector Machine was determined to be optimal, with the option of training using solely the VH polarization or a combination of the VH and VV polarizations. WASARD was able to identify water in the target area with a correlation of 97% with WOFS. Sentinel-1, Open Data Cube, Earth Observations, Machine Learning, Water Detection 1. INTRODUCTION Water classification is an important function of Earth imaging satellites, as accurate remote classification of land and water can assist in land use analysis, flood prediction, climate change research, as well as a variety of agricultural applications [2]. The ability to identify bodies of water remotely via satellite is immensely cheaper than contracting surveys of the areas in question, meaning that an application that can accurately use satellite data towards this function can make valuable information available to nations which would not be able to afford it otherwise. Highly reliable applications for the remote detection of water currently exist for use with optical satellite data such as that provided by LANDSAT. One such application, Geoscience Australia’s Water Observations from Space (WOFS) has already been ported for use with the Open Data Cube [6]. However, water detection using optical data from Landsat is constrained by its relatively long revisit cycle of 16 days [5], and water detection using any optical data is constrained in that it lacks the ability to make accurate classifications through cloud cover [2]. The alternative solution which solves these problems is water detection using SAR data, which images the Earth using cloud-penetrating microwaves. Because of its advantages over optical data, much research has been done into water detection using SAR data. Traditionally, this has been done using the thresholding method, which involves picking a polarization band and labeling all pixels for which this band’s value is below a certain threshold as containing water. The thresholding method works since water tends to return a much lower backscatter value to the satellite than land [1]. However, this method can be flawed since estimating the proper threshold is often imprecise, complicated, and labor intensive for the end user. Thresholding also tends to use data from only one SAR polarization, when a combination of polarizations can provide insight into whether water is present. [2] In order to alleviate these problems, this paper presents an application for the Open Data Cube to detect water from SAR data using support vector machine (SVM) classification. 2. PLATFORM WASARD is an application for the Open Data Cube, a mechanism which provides a simple yet efficient means of ingesting, storing, and retrieving remote sensing data. Data can be ingested and made analysis ready according to whatever specifications the researcher chooses, and easily resampled to artificially alter a scene’s resolution. Currently WASARD supports water detection on scenes from ESA’s Sentinel-1 and JAXA’s ALOS. When testing WASARD, Sentinel-1 was most commonly used due to its relatively high spatial resolution and its rapid 6 day revisit cycle [5]. With minor alterations to the application's code, however, it could support data from other satellites. 3. METHODOLOGY Using supervised classification, WASARD compares SAR data to a dataset pre-classified by WOFS in order to train an SVM classifier. This classifier is then used to detect water in other SAR scenes outside the training set. Accuracy was measured according to the following metrics:  Precision: a measure of what percentage of the points WASARD labels as water are truly water  Recall: a measure of what percentage of the total water cover WASARD was able to identify.  F1 Score: a harmonic average of the precision and recall scores Both precision and recall are calculated at the end of the training phase, when the trained classifier is compared to a testing dataset. Because the WOFS algorithm’s classifications are used as the truth values when training a WASARD classifier, when precision and recall are mentioned in this paper, they are always with respect to the values produced by WOFS on a similar scene of Landsat data, which themselves have a classification accuracy of 97% [6]. Visual representations of water identified by WASARD in this paper were produced using the function wasard_plot(), which is included in WASARD. 3.1 Algorithm Selection The machine learning model used by WASARD is the Linear Support Vector Machine (SVM). This model uses a supervised learning algorithm to develop a classifier, meaning it creates a vector which can be multiplied by the vector formed by the relevant data bands to determine whether a pixel in a SAR scene contains water. This classifier is trained by comparing data points from selected bands in a SAR scene to their respective labels, which in this case are “water” or “not water” as given by the WOFS algorithm. The SVM was selected over the Random Forest model, which outperformed the SVM in training speed, but had a greater classification time and lower accuracy, and the Multilayer Perceptron Artificial Neural Network, which had a slightly higher average accuracy than the SVM, but much greater training and classification times. Figure 1: Visual representation of the SVM Classifier. Each white point represents a pixel in a SAR scene. In Figure 1, the diagonal line separating pixels determined to be water from those determined not to be water represents the actual classification vector produced by the SVM. It is worth noting that once the model has been trained, classification of pixels is done in a similar manner as in the thresholding method. This is especially true if only one band was used to train the model. 3.1 Feature Selection Sentinel-1 collects data from two bands: the Vertical/Vertical polarization (VV) and the Vertical/Horizontal polarization (VH). When 100 SVM classifiers were created for each polarization individually, and for the combination of the two, the following results were achieved: Figure 2: Accuracy of classifiers trained using different polarization bands. Precision and Recall were measured with respect to the values produced by WOFS. Figure 2 demonstrates that using both the VV and VH bands trades slightly lower recall for significantly greater precision when compared with the VH band alone, and that using the VV band alone is inferior in both metrics. WASARD therefore defaults to using both the VV and VH bands, and includes the option to use solely the VH band. The VV polarization’s lower precision compared to the VH polarization is in contrast to results from previous research and may merit further analysis [4]. 3.2 Training a Classifier The steps in training a classifier with WASARD are 1. Selecting two scenes (one SAR, one optical) with the same spatial extents, and acquired close to each other in time, with a preference that the scenes are taken on the same day. 2. Using the WOFS algorithm to produce an array of the detected water in the scene of optical data, to be used as the labels during supervised learning 3. Data points from the selected bands from the SAR acquisition are bundled together into an array with the corresponding labels gathered from WOFS. A random sample with an equal number of points labeled “Water” and “Not Water” is selected to be partitioned into a training and a testing dataset 4. Using Scikit-Learn’s LinearSVC object, the training dataset is used to produce a classifier, which is then tested against the testing dataset to determine its precision and recall The result is a wasard_classifier object, which has the following attributes: 1. f1, recall, and precision: 3 metrics used to determine the classifier’s accuracy 2. Coefficient: Vector which the SVM uses to make its predictions. The classifier detects water when the dot product of the coefficient and the vector formed by the SAR bands is positive 3. Save(): allows a user to save a classifier to the disk in order to use it without retraining 4. wasard_classify(): Classifies an entire xarray of SAR data using the SVM classifier All of the above steps are performed automatically when the user creates a wasard_classifier object. 3.3 Classifying a Dataset Once the classifier has been created, it can be used to detect water in an xarray of SAR data using wasard_classify(). By taking the dot product of the classifier’s coefficients and the vector formed by the selected bands of SAR data, an array of predictions is constructed. A classifier can effectively be used on the same spatial extents as the ones where it was trained, or on any area with a similar landscape. While

Kreiser, Zachary↗

Multiple-Instance Regression with Structured Data

We present a multiple-instance regression algorithm that models internal bag structure to identify the items most relevant to the bag labels. Multiple-instance regression (MIR) operates on a set of bags with real-valued labels, each containing a set of unlabeled items, in which the relevance of each item to its bag label is unknown. The goal is to predict the labels of new bags from their contents. Unlike previous MIR methods, MI-ClusterRegress can operate on bags that are structured in that they contain items drawn from a number of distinct (but unknown) distributions. MI-ClusterRegress simultaneously learns a model of the bag's internal structure, the relevance of each item, and a regression model that accurately predicts labels for new bags. We evaluated this approach on the challenging MIR problem of crop yield prediction from remote sensing data. MI-ClusterRegress provided predictions that were more accurate than those obtained with non-multiple-instance approaches or MIR methods that do not model the bag structure.

learning↗

Foundation AI Models for Science

Foundation Models (FM) are AI models that are designed to replace a task or an application specific model. These FM can be applied to many different downstream applications. These FM are trained using self supervised techniques and can be built on any type of sequence data. The use of self supervised learning removes the hurdle for developing a large labeled dataset for training. Most FM use transformer architecture utilizes the notion of self attention which allows the network to model the influence of distant data points to each other both in space and time. The FM models exhibit emergent properties that are induced from the data. FM can be an important tool for science. The scale of these models results in better performance for different downstream applications and these applications show better accuracy over models built from scratch. FM drastically reduces the cost of entry to build different downstream applications both in time and effort. FM for selected science datasets such as optical satellite data, can accelerate applications ranging from data quality monitoring, feature detection and prediction. FM can make it easier to infuse AI into scientific research by removing the training data bottleneck and increasing the use of science data.

Manil Maskey↗

Evaluation of ceramic packed-rod regenerator matrices

An extensive evaluation of a modified cryocooler with various regenerator matrices is reported. The matrices examined are 0.015 in. diam. Pb spheres and 0.008, 0.015, and 0.030 in. diam. rods of a 0.2% SnCl2 doped ceramic labelled LS-8A. Specific heat and thermal conductivity data on these rod materials are also reported. The chronic pulverization/dusting problem common to Pb spheres was investigated. During a 1000 hr life test with 0.0008 in. diam. rods there was no degradation of the refrigerator performance, and a subsequent examination of the rods themselves revealed no evidence of breakage or pulverization. The load temperature characteristics for the rod packed regenerators were inferior to that for the Pb spheres, the effect being to shift the Pb spheres load curve up in temperature. This temperature shift was 5.0, 7.4, and 11.6K for the 0.0008, 0.015, and 0.030 in. diam. rods, respectively.

Lawless, W. N.↗

Description of Data Archiving Activities

Data restoration and archiving activities for this project have resulted in the restoration of 100% of the original Mariner 9 raw data set as well as many of the secondary analysis data sets. These data sets have been submitted to the Planetary Data System (PDS) Atmospheric Node, long with their PDS labels and descriptive metadata.

Simmons, K. E.↗

Characterizing the Oxidizing Properties of Mars' Polar Regions

This project had two primary goals. The first was to restore and archive the Ultraviolet Spectrometer (UVS) data from the 1971 Mariner 9 (MM71) mission to Mars. The second was to use this revised data set to analyze data of Mars' polar regions to look for and map out the ozone (03) and hydrogen peroxide (H2O2,) features. Data restoration and archiving activities for this project have resulted in the restoration of 100% of the original Mariner 9 raw data set as well as many of the secondary analysis data sets. These data sets have been submitted to the Planetary Data System (PDS) Atmospheric Node, long with their PDS labels and descriptive metadata.

Hendrix, Amanda↗

Automatic labeling and characterization of objects using artificial neural networks

Existing NASA supported scientific data bases are usually developed, managed and populated in a tedious, error prone and self-limiting way in terms of what can be described in a relational Data Base Management System (DBMS). The next generation Earth remote sensing platforms, i.e., Earth Observation System, (EOS), will be capable of generating data at a rate of over 300 Mbs per second from a suite of instruments designed for different applications. What is needed is an innovative approach that creates object-oriented databases that segment, characterize, catalog and are manageable in a domain-specific context and whose contents are available interactively and in near-real-time to the user community. Described here is work in progress that utilizes an artificial neural net approach to characterize satellite imagery of undefined objects into high-level data objects. The characterized data is then dynamically allocated to an object-oriented data base where it can be reviewed and assessed by a user. The definition, development, and evolution of the overall data system model are steps in the creation of an application-driven knowledge-based scientific information system.

Campbell, William J.↗

Remote sensing data processing - Two years ago, today, and two years from today

Certain technical problems arising in the recent past (1975) in the field of the processing of remote sensing data are reviewed including approaches to the analysis of Landsat MSS data and technical difficulties which must be overcome to achieve operational data processing. The current status of remote sensing data processing is then examined with emphasis on such current technical issues as training selection and labeling, sampling schemes and classification and mensuration. Hardware projections are made for the near future (1979) relative to the development of remote sensing data processing.

Holmes, Q. A.↗

User's guide for the Solar Backscattered Ultraviolet (SBUV) and the Total Ozone Mapping Spectrometer (TOMS) RUT-S and RUT-T data sets: October 31, 1978 to November 1, 1980

Raw data from the Solar Backscattered Ultrviolet/Total Ozone Mapping Spectrometer (SBUV/TOMS) Nimbus 7 operation are available on computer tape. These data are contained on two separate sets of RUTs (Raw Units Tapes) for SBUV and TOMS, labelled RUT-S and RUT-T respectively. The RUT-S and RUT-T tapes contain uncalibrated radiance and irradiance data, housekeeping data, wavelength and electronic calibration data, instrument field-of-view location and solar ephemeris information. These tapes also contain colocated cloud, terrain pressure and snow/ice thickness data, each derived from an independent source. The "RUT User's Guide" describes the SBUV and TOMS experiments, the instrument calibration and performance, operating schedules, and data coverage, and provides an assessment of RUT-S and -T data quality. It also provides detailed information on the data available on the computer tapes.

Fleig, A. J.↗

Optical imaging and long-slit spectroscopy of Markarian galaxies with multiple nuclei. I - Basic data

Optical CCD images and long-slit spectroscopic data are presented for over 100 Markarian (UV-excess) galaxies reported in early studies to possess multiple optical nuclei or extreme morphological peculiarities suggestive of galaxy collisions and mergers. Stacked broad-band images are presented with histogram equalization in order to show simultaneously the nuclei and features at very low surface-brightness levels. Morphological properties, luminosities and colors of the integral systems are given. Photometric and image properties of over 200 individual nuclei and giant H II regions have been measured with respect to the local backgrounds in the galaxies using an objective image finding algorithm. Labeled contour plots identify the measured subcomponents. Two-dimensional spectral data are presented, in addition to intensity profiles along the slit in the light of H-alpha + forbidden N II emission lines and adjacent continuum. Nuclear emission-line measurements, reddening estimates, monochromatic continuum magnitudes, and colors are given.

Mazzarella, Joseph M.↗

Evidence of hyperbolic cosmic dust particles.

Cosmic dust data from the helicentric Pioneers 8 and 9 have been gathered for more than 7 years. A review and detailed study of these data are given and show that events which were previously labeled solar disturbance events and assumed to be noise generated by solar effects are, logically, true cosmic dust impact events. They are shown as an extension of the range of particle parameters exhibited by the time-of-flight measurements. The effects of accepting the sun-oriented events as authentic impact events are discussed.

Berg, O. E.↗

Clustering algorithm evaluation and the development of a replacement for procedure 1

An efficient procedure which clusters data using a completely unsupervised clustering algorithm and then uses labeled pixels to label the resulting clusters or perform a stratified estimate using the clusters as strata is developed. Three clustering algorithms, CLASSY, AMOEBA, and ISOCLS, are compared for efficiency. Three stratified estimation schemes and three labeling schemes are also considered and compared.

Lennington, R. K.↗

LACIE analyst interpretation keys

Two interpretation aids, 'The Image Analysis Guide for Wheat/Small Grains Inventories' and 'The United States and Canadian Great Plains Regional Keys', were developed during LACIE phase 2 and implemented during phase 3 in order to provide analysts with a better understanding of the expected ranges in color variation of signatures for individual biostages and of the temporal sequences of LANDSAT signatures. The keys were tested using operational LACIE data, and the results demonstrate that their use provides improved labeling accuracy in all analyst experience groupings, in all geographic areas within the U.S. Great Plains, and during all periods of crop development.

Baron, J. G.↗

A programmed labeling approach to image interpretation

Manual labeling techniques require the analyst-interpreter to use not only production film converter products but also agricultural and meteorological data and spectral aids in an integrated, judgmental fashion. To control an anticipated high variance in these techniques, a semiautomatic labeling technology was developed. The product of this technology is label identification from statistical tabulation (LIST) which operates from a discriminant basis and has the ability to measure the reliability of the label and to introduce an arbitrary bias. The development of LIST and its properties are described. Numerical results of an application are included and the evaluation of LIST is discussed.

Pore, M. D.↗