Search NASA⌕ Search

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Foundation AI Models for Science

Foundation Models (FM) are AI models that are designed to replace a task or an application specific model. These FM can be applied to many different downstream applications. These FM are trained using self supervised techniques and can be built on any type of sequence data. The use of self supervised learning removes the hurdle for developing a large labeled dataset for training. Most FM use transformer architecture utilizes the notion of self attention which allows the network to model the influence of distant data points to each other both in space and time. The FM models exhibit emergent properties that are induced from the data. FM can be an important tool for science. The scale of these models results in better performance for different downstream applications and these applications show better accuracy over models built from scratch. FM drastically reduces the cost of entry to build different downstream applications both in time and effort. FM for selected science datasets such as optical satellite data, can accelerate applications ranging from data quality monitoring, feature detection and prediction. FM can make it easier to infuse AI into scientific research by removing the training data bottleneck and increasing the use of science data.

Manil Maskey↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Evaluation of ceramic packed-rod regenerator matrices

An extensive evaluation of a modified cryocooler with various regenerator matrices is reported. The matrices examined are 0.015 in. diam. Pb spheres and 0.008, 0.015, and 0.030 in. diam. rods of a 0.2% SnCl2 doped ceramic labelled LS-8A. Specific heat and thermal conductivity data on these rod materials are also reported. The chronic pulverization/dusting problem common to Pb spheres was investigated. During a 1000 hr life test with 0.0008 in. diam. rods there was no degradation of the refrigerator performance, and a subsequent examination of the rods themselves revealed no evidence of breakage or pulverization. The load temperature characteristics for the rod packed regenerators were inferior to that for the Pb spheres, the effect being to shift the Pb spheres load curve up in temperature. This temperature shift was 5.0, 7.4, and 11.6K for the 0.0008, 0.015, and 0.030 in. diam. rods, respectively.

Lawless, W. N.↗

Description of Data Archiving Activities

Data restoration and archiving activities for this project have resulted in the restoration of 100% of the original Mariner 9 raw data set as well as many of the secondary analysis data sets. These data sets have been submitted to the Planetary Data System (PDS) Atmospheric Node, long with their PDS labels and descriptive metadata.

Simmons, K. E.↗

Characterizing the Oxidizing Properties of Mars' Polar Regions

This project had two primary goals. The first was to restore and archive the Ultraviolet Spectrometer (UVS) data from the 1971 Mariner 9 (MM71) mission to Mars. The second was to use this revised data set to analyze data of Mars' polar regions to look for and map out the ozone (03) and hydrogen peroxide (H2O2,) features. Data restoration and archiving activities for this project have resulted in the restoration of 100% of the original Mariner 9 raw data set as well as many of the secondary analysis data sets. These data sets have been submitted to the Planetary Data System (PDS) Atmospheric Node, long with their PDS labels and descriptive metadata.

Hendrix, Amanda↗

Automatic labeling and characterization of objects using artificial neural networks

Existing NASA supported scientific data bases are usually developed, managed and populated in a tedious, error prone and self-limiting way in terms of what can be described in a relational Data Base Management System (DBMS). The next generation Earth remote sensing platforms, i.e., Earth Observation System, (EOS), will be capable of generating data at a rate of over 300 Mbs per second from a suite of instruments designed for different applications. What is needed is an innovative approach that creates object-oriented databases that segment, characterize, catalog and are manageable in a domain-specific context and whose contents are available interactively and in near-real-time to the user community. Described here is work in progress that utilizes an artificial neural net approach to characterize satellite imagery of undefined objects into high-level data objects. The characterized data is then dynamically allocated to an object-oriented data base where it can be reviewed and assessed by a user. The definition, development, and evolution of the overall data system model are steps in the creation of an application-driven knowledge-based scientific information system.

Campbell, William J.↗

Remote sensing data processing - Two years ago, today, and two years from today

Certain technical problems arising in the recent past (1975) in the field of the processing of remote sensing data are reviewed including approaches to the analysis of Landsat MSS data and technical difficulties which must be overcome to achieve operational data processing. The current status of remote sensing data processing is then examined with emphasis on such current technical issues as training selection and labeling, sampling schemes and classification and mensuration. Hardware projections are made for the near future (1979) relative to the development of remote sensing data processing.

Holmes, Q. A.↗

User's guide for the Solar Backscattered Ultraviolet (SBUV) and the Total Ozone Mapping Spectrometer (TOMS) RUT-S and RUT-T data sets: October 31, 1978 to November 1, 1980

Raw data from the Solar Backscattered Ultrviolet/Total Ozone Mapping Spectrometer (SBUV/TOMS) Nimbus 7 operation are available on computer tape. These data are contained on two separate sets of RUTs (Raw Units Tapes) for SBUV and TOMS, labelled RUT-S and RUT-T respectively. The RUT-S and RUT-T tapes contain uncalibrated radiance and irradiance data, housekeeping data, wavelength and electronic calibration data, instrument field-of-view location and solar ephemeris information. These tapes also contain colocated cloud, terrain pressure and snow/ice thickness data, each derived from an independent source. The "RUT User's Guide" describes the SBUV and TOMS experiments, the instrument calibration and performance, operating schedules, and data coverage, and provides an assessment of RUT-S and -T data quality. It also provides detailed information on the data available on the computer tapes.

Fleig, A. J.↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

Optical imaging and long-slit spectroscopy of Markarian galaxies with multiple nuclei. I - Basic data

Optical CCD images and long-slit spectroscopic data are presented for over 100 Markarian (UV-excess) galaxies reported in early studies to possess multiple optical nuclei or extreme morphological peculiarities suggestive of galaxy collisions and mergers. Stacked broad-band images are presented with histogram equalization in order to show simultaneously the nuclei and features at very low surface-brightness levels. Morphological properties, luminosities and colors of the integral systems are given. Photometric and image properties of over 200 individual nuclei and giant H II regions have been measured with respect to the local backgrounds in the galaxies using an objective image finding algorithm. Labeled contour plots identify the measured subcomponents. Two-dimensional spectral data are presented, in addition to intensity profiles along the slit in the light of H-alpha + forbidden N II emission lines and adjacent continuum. Nuclear emission-line measurements, reddening estimates, monochromatic continuum magnitudes, and colors are given.

Mazzarella, Joseph M.↗

On the Prospect of Chemically Transferable Coarse-Grained Electronic Models for Soft Materials

Electronic coarse-graining (ECG) methods predict quantum-mechanical electronic properties directly from coarse-grained (CG) molecular configurations, enabling electronic predictions at mesoscale length scales. Here, we present a diagnostic assessment of the feasibility of chemically transferable ECG models across a broad polymer-relevant chemical space using all-atom, united-atom, and Martini-scale representations. While high-resolution ECG models achieve near-quantitative accuracy, we show that chemically transferable ECG at the Martini resolution fails because the CG force field does not sample the same configurational distribution of local molecular structure as that underlying the DFT-parameterized ECG model. We demonstrate that our proposed Element-Count-Label (ECL) representation, which augments Martini beads with explicit stoichiometric data, significantly improves chemical generalization across diverse polymer chemistries. However, we find that even with improved chemical resolution, the model cannot recover electronic property distributions that are absent from the configurational space sampled by the CG force field. These results demonstrate that chemically transferable ECG requires future Martini-like force fields to explicitly preserve quantum chemistry–compatible local molecular structure in addition to thermodynamic and structural fidelity.

Kidder, Katherine M [Department of Chemistry; Univ↗

Evidence of hyperbolic cosmic dust particles.

Cosmic dust data from the helicentric Pioneers 8 and 9 have been gathered for more than 7 years. A review and detailed study of these data are given and show that events which were previously labeled solar disturbance events and assumed to be noise generated by solar effects are, logically, true cosmic dust impact events. They are shown as an extension of the range of particle parameters exhibited by the time-of-flight measurements. The effects of accepting the sun-oriented events as authentic impact events are discussed.

Berg, O. E.↗

Clustering algorithm evaluation and the development of a replacement for procedure 1

An efficient procedure which clusters data using a completely unsupervised clustering algorithm and then uses labeled pixels to label the resulting clusters or perform a stratified estimate using the clusters as strata is developed. Three clustering algorithms, CLASSY, AMOEBA, and ISOCLS, are compared for efficiency. Three stratified estimation schemes and three labeling schemes are also considered and compared.

Lennington, R. K.↗

LACIE analyst interpretation keys

Two interpretation aids, 'The Image Analysis Guide for Wheat/Small Grains Inventories' and 'The United States and Canadian Great Plains Regional Keys', were developed during LACIE phase 2 and implemented during phase 3 in order to provide analysts with a better understanding of the expected ranges in color variation of signatures for individual biostages and of the temporal sequences of LANDSAT signatures. The keys were tested using operational LACIE data, and the results demonstrate that their use provides improved labeling accuracy in all analyst experience groupings, in all geographic areas within the U.S. Great Plains, and during all periods of crop development.

Baron, J. G.↗

A programmed labeling approach to image interpretation

Manual labeling techniques require the analyst-interpreter to use not only production film converter products but also agricultural and meteorological data and spectral aids in an integrated, judgmental fashion. To control an anticipated high variance in these techniques, a semiautomatic labeling technology was developed. The product of this technology is label identification from statistical tabulation (LIST) which operates from a discriminant basis and has the ability to measure the reliability of the label and to introduce an arbitrary bias. The development of LIST and its properties are described. Numerical results of an application are included and the evaluation of LIST is discussed.

Pore, M. D.↗

Determining Stellar Elemental Abundances from DESI Spectra with the Data-driven Payne

Abstract Stellar abundances for a large number of stars provide key information for the study of Galactic formation history. Large spectroscopic surveys such as the Dark Energy Spectroscopic Instrument (DESI) and LAMOST take median-to-low-resolution (R≲ 5000) spectra in the full optical wavelength range for millions of stars. However, the line-blending effect in these spectra causes great challenges for elemental abundance determination. Here we employDD-Payne, a data-driven method regularized by differential spectra from stellar physical models, to the DESI early data release spectra for stellar abundance determination. Our implementation delivers 15 labels, including effective temperatureT eff , surface gravity log g , microturbulence velocityv mic , and the abundances for 12 individual elements, namely C, N, O, Mg, Al, Si, Ca, Ti, Cr, Mn, Fe, and Ni. Given a spectral signal-to-noise ratio of 100 per pixel, the internal precisions of the label estimates are about 20 K forT eff , 0.05 dex for log g , and 0.05 dex for most elemental abundances. These results agree with the theoretical limits from the Crámer–Rao bound calculation within a factor of 2. The majority of the accreted halo stars contributed by the Gaia–Enceladus–Sausage are discernible from the disk and in situ halo populations in the resultant [Mg/Fe]–[Fe/H] and [Al/Fe]–[Fe/H] abundance spaces. We also provide distance and orbital parameters for the sample stars, which spread over a distance out to ∼100 kpc. The DESI sample has a significantly higher fraction of distant (or metal-poor) stars than the other existing spectroscopic surveys, making it a powerful data set for studying the Galactic outskirts. The catalog is publicly available.

Astronomy & Astrophysics↗

Single-class classification

Often, when classifying multispectral data, only one class or crop is of interest, such as wheat in the Large Area Crop Inventory Experiment (LACIE). Usual procedures for designing a Bayes classifier require that labeled training samples and therefore ground truth be available for the 'class of interest' plus all confusion classes defined by the multispectral data. This paper will consider the problem of designing a two-class Bayes classifier which will classify data into the 'class of interest' or the 'other' classes but will require only labeled training samples from the 'class of interest' to design the classifier. Thus, this classifier minimizes the need for ground truth. For these reasons, the classifier is referred to as a single-class classifier. A procedure for evaluating the overall performance of the single-class classifier in terms of the probability of error will be discussed.

Minter, T. C.↗