Search NASA⌕ Search

SEARCH · Search NASA

Results for “Labeled Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Semiempirical Estimate of Aircraft Wing Weight

Computational method estimates weight of aircraft wings from theoretical relationships and empirical data. Permits comparison of alternative materials, methods of construction, and design philosophies. Method used to make tradeoffs in preliminary design phases on basis of simple input data and for more accurate calculations in later phases when more data are available.

York, P.↗

The diageotropica mutant of tomato lacks high specific activity auxin binding sites

Tomato plants homozygous for the diageotropica (dgt) mutation exhibit morphological and physiological abnormalities which suggest that they are unable to respond to the plant growth hormone auxin (indole-3-acetic acid). The photoaffinity auxin analog [3H]5N3-IAA specifically labels a polypeptide doublet of 40 and 42 kilodaltons in membrane preparations from stems of the parental variety, VFN8, but not from stems of plants containing the dgt mutation. In roots of the mutant plants, however, labeling is indistinguishable from that in VFN8. These data suggest that the two polypeptides are part of a physiologically important auxin receptor system, which is altered in a tissue-specific manner in the mutant.

NASA Discipline Number 29-20↗

Data integration in multi-sensor based robotic workstations

A relaxation labeling algorithm is developed. The major advantage of this algorithm over the existing ones is that the mathematic operation is simplified. The simplification eases the analysis of the convergence properties. Both the theoretical and application aspects of the proposed algorithm are investigated. The local convergence properties of a labeling process with n labels and m labels are established. The investigation of the interaction among the nodes in a multinode labeling process reveals some insight into the mathematical issues involved in the relaxation operations.

Chen, Qin↗

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES↗

Automated cloud classification with a fuzzy logic expert system

An unresolved problem in current cloud retrieval algorithms concerns the analysis of scenes containing overlapping cloud layers. Cloud parameterizations are very important both in global climate models and in studies of the Earth's radiation budget. Most cloud retrieval schemes, such as the bispectral method used by the International Satellite Cloud Climatology Project (ISCCP), have no way of determining whether overlapping cloud layers exist in any group of satellite pixels. One promising method uses fuzzy logic to determine whether mixed cloud and/or surface types exist within a group of pixels, such as cirrus, land, and water, or cirrus and stratus. When two or more class types are present, fuzzy logic uses membership values to assign the group of pixels partially to the different class types. The strength of fuzzy logic lies in its ability to work with patterns that may include more than one class, facilitating greater information extraction from satellite radiometric data. The development of the fuzzy logic rule-based expert system involves training the fuzzy classifier with spectral and textural features calculated from accurately labeled 32x32 regions of Advanced Very High Resolution Radiometer (AVHRR) 1.1-km data. The spectral data consists of AVHRR channels 1 (0.55-0.68 mu m), 2 (0.725-1.1 mu m), 3 (3.55-3.93 mu m), 4 (10.5-11.5 mu m), and 5 (11.5-12.5 mu m), which include visible, near-infrared, and infrared window regions. The textural features are based on the gray level difference vector (GLDV) method. A sophisticated new interactive visual image Classification System (IVICS) is used to label samples chosen from scenes collected during the FIRE IFO II. The training samples are chosen from predefined classes, chosen to be ocean, land, unbroken stratiform, broken stratiform, and cirrus. The November 28, 1991 NOAA overpasses contain complex multilevel cloud situations ideal for training and validating the fuzzy logic expert system.

Tovinkere, Vasanth↗

The Radar Image Generation (RIG) model

RIG is a modeling system which creates synthetic aperture radar (SAR) and inverse SAR images from 3-D faceted data bases. RIG is based on a physical optics model and includes the effects of multiple reflections. Both conducting and dielectric surfaces can be modeled; each surface is labeled with a material code which is an index into a data base of electromagnetic properties. The inputs to the program include the radar processing parameters, the target orientation, the sensor velocity, and (for inverse SAR) the target angle rates. The current version of RIG can be run on any workstation, however, it is not a real-time model. We are considering several approaches to enable the program to generate realtime radar imagery. In addition to its image generation function, RIG can also generate radar cross-section (RCS) plots as well as range and doppler radar return profiles.

Stenger, Anthony J.↗

Is tokenization needed for masked particle modeling?

In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics. In MPM, a model is trained to recover the missing elements of a set, a learning objective that requires no labels and can be applied directly to experimental data. We achieve significant performance improvements over previous work on MPM by addressing inefficiencies in the implementation and incorporating a more powerful decoder. We compare several pre-training tasks and introduce new reconstruction methods that utilize conditional generative models without data tokenization or discretization. We show that these new methods outperform the tokenized learning objective from the original MPM on a new test bed for foundation models for jets, which includes using a wide variety of downstream tasks relevant to jet physics, such as classification, secondary vertex finding, and track identification.

conditional generative models↗

Program Manipulates Plots For Effective Display

Windowed Observation of Relative Motion (WORM) computer program primarily intended for generation of simple X-Y plots from data created by other programs. Enables user to label, zoom, and change scales of various plots. Three-dimensional contour and line plots provided. Written in PASCAL.

Bauer, F.↗

Automatic Detection and Classification of Aurora in THEMIS All‐Sky Images

We report a novel machine-learning algorithm for automatically detecting and classifying aurora in all–sky images (ASI) that is largely trained without requiring ground–truth labels. By including a small number of labeled images, we are able to automatically label all of the approximately 700 million images in the Time History of Events and Macroscale Interactions during Substorms (THEMIS) ASI data set from 2008 to 2022. We use a two–stage approach. In the first stage, we adapt the Simple framework for Contrastive Learning of Representations (SimCLR) algorithm to learn latent representations of THEMIS all–sky images. We then finetune a classifier network on the latent representations our model learns of the manually labeled Oslo aurora THEMIS (OATH) data set. We demonstrate that this two–stage approach achieves excellent classification results on data for which there is no current ML classification benchmark. The outcome of this work will facilitate efficient information retrieval for researchers interested in specific categories of aurora and will enable large scale statistical studies and machine learning analyses of THEMIS all–sky images that have not previously been possible. To demonstrate possible ways to utilize this database, we performed a statistical analysis of the occurrence rates of auroral labels with respect to solar wind parameters, interplanetary magnetic field vector, and geomagnetic indices. We further investigate the occurrence rates of auroral phenomena in the annotated data set and their geoeffectiveness by utilizing the co–located THEMIS ground magnetometer data set.

Jeremiah W Johnson↗

Application of modified VICAR/IBIS GIS to analysis of July 1991 Flevoland AIRSAR data

Three overflights of the Flevoland calibration/agricultural site were made by the JPL Airborne Synthetic Aperture Radar (AIRSAR) on 3, 12, and 28 July 1991 as part of MAC-Europe '92. A polygon map was generated at TNO-FEL which overlayed the slant range projected July 3 data set. Each polygon was identified by a sequence of points and a crop label. The polygon map was composed of 452 uniquely identified polygons and 15 different crop types. Analysis of the data was done using our modified Video Image Communication and Retrieval/Image Based Information System Geographic Information System (VICAR/IBIS GIS). This GIS is an extension of the VICAR/IBIS GIS first developed by Bryant in the 1970's which is itself an extension of the VICAR image processing system also developed at JPL.

Norikane, L.↗

Information science team

Concerns are expressed about the data handling aspects of system design and about enabling technology for data handling and data analysis. The status, contributing factors, critical issues, and recommendations for investigations are listed for data handling, rectification and registration, and information extraction. Potential supports to individual P.I., research tasks, systematic data system design, and to system operation. The need for an airborne spectrometer class instrument for fundamental research in high spectral and spatial resolution is indicated. Geographic information system formatting and labelling techniques, very large scale integration, and methods for providing multitype data sets must also be developed.

Billingsley, F.↗

LLMs and GenAI Tools to Depict Contributions of Human Systems to Spaceflight Tasks Execution

Recent advancements in Artificial Intelligence and Machine Learning (AI/ML) technologies, particularly Large Language Models (LLMs) capable of sophisticated syntax analysis, offer substantial potential in automating complex processes, thereby saving time and human resources. This study explores the development of an LLM-driven model designed to analyze and categorize a diverse set of Mars mission tasks into 18 predefined Human System Task Categories (HSTCs) based on their textual descriptions. As part of developing the Crew Health and Performance – Probabilistic Risk Assessment (CHP-PRA projects Performance Risk Model (PRisM) proof-of-concept, we established a framework to project performance scores from small-scale tests onto a preliminary list of Mars tasks. The foundation of our model was a comprehensive spreadsheet populated by NASA experts and clinicians, which detailed each Mars task alongside binary indicators of HSTC involvement. This dataset enabled the initial application of supervised ML, training and testing on existing HSTC labels. The HSTCs were originally defined from a medical system perspective, focusing on task impairments due to deteriorated human health. To expand our model's scope to include categories impacting performance, we face the challenge of generating binary labels (0 or 1) for new categories without pre-existing data. We address this by employing Generative AI (GenAI) software to determine whether a given task involved a new category by asking, "Does task A involve using category B?" We validate our approach by comparing the GenAI's binary classifications with the expert-provided labels for existing HSTCs. Notably, we utilize Ollama [4], a locally hosted GenAI tool that does not require cloud access, thus safeguarding NASA's proprietary data from unauthorized exposure. This study demonstrates the feasibility of leveraging cutting-edge AI tools to advance research, paving the way for automation and rapid decision-making in space exploration.

Mona Matar↗

A self-documenting source-independent data format for computer processing of tensor time series

The UCLA Space Science Group has developed a fixed format intermediate data set called a block data set, which is designed to hold multiple segments of multicomponent sampled data series. The format is sufficiently general so that tensor functions of one or more independent variables can be stored in the form of virtual data. This makes it possible for the unit data records of the block data set to be arrays of a single dependent variable rather than discrete samples. The format is self-documenting with parameter, label and header records completely characterizing the contents of the file. The block data set has been applied to the filing of satellite data (of ATS-6 among others).

Mcpherron, R. L.↗

Mars Terrain Segmentation with Less Labels

Planetary rover systems need to perform terrain segmentation to identify drivable areas as well as identify specific types of soil for sample collection. The latest Martian terrain segmentation methods rely on supervised learning which is very data hungry and difficult to train where only a small number of labeled samples are available. Moreover, the semantic classes are defined differently for different applications (e.g., rover traversal vs. geological) and as a result the network has to be trained from scratch each time, which is an inefficient use of resources. This research proposes a semi-supervised learning framework for Mars terrain segmentation where a deep segmentation network trained in an unsupervised manner on unlabeled images is transferred to the task of terrain segmentation trained on few labeled images. The network incorporates a backbone module which is trained using a contrastive loss function and an output atrous convolution module which is trained using a pixel-wise cross-entropy loss function. Evaluation results using the metric of segmentation accuracy show that the proposed method with contrastive pre-training outperforms plain supervised learning by 2%-10%. Moreover, the proposed model is able to achieve a segmentation accuracy of 91.1% using only 161 training images (1% of the original dataset) compared to 81.9% with plain supervised learning.

Wilson, Brian D↗

A Morphological Model to Separate Resolved–Unresolved Sources in the DESI Legacy Surveys: Application in the LS4 Alert Stream

Separating resolved and unresolved sources in large imaging surveys is a fundamental step to enable downstream science, such as searching for extragalactic transients in wide-field time-domain surveys. Here we present our method to effectively separate point sources from the resolved, extended sources in the Dark Energy Spectroscopic Instrument (DESI) Legacy Surveys (LS). We develop a supervised machine learning model based on the Gradient Boosting algorithm XGBoost. The features input to the model are purely morphological and are derived from the tabulated LS data products. We train the model using ∼2 × 10 5 LS sources in the COSMOS field with HST morphological labels and evaluate the model performance on LS sources with spectroscopic classification from the DESI Data Release 1 (∼2 × 10 7 objects) and the Sloan Digital Sky Survey Data Release 17 (∼3 × 10 6 objects), as well as on ∼2 × 10 8 Gaia stars. A significant fraction of LS sources are not observed in every LS filter, and we therefore build a “Hybrid” model as a linear combination of two XGBoost models, each containing features combining aperture flux measurements from the “blue” (gr) and “red” (iz) filters. The Hybrid model shows a reasonable balance between sensitivity and robustness, and achieves higher accuracy and flexibility compared to the LS morphological typing. With the Hybrid model, we provide classification scores for ∼3 × 10 9 LS sources, making this the largest ever machine learning catalog separating resolved and unresolved sources. The catalog has been incorporated into the real-time pipeline of the La Silla Schmidt Southern Survey (LS4), enabling the identification of extragalactic transients within the LS4 alert stream.

astrostatistics↗

Cleaned 5-Minute Resolution Air Quality and Meteorological Data from Nine TCEQ CAMS Sites in Houston, Texas (Nov 2021 – Oct 2022)

These data encompass 5-minute air monitoring and meteorological observations collected in the greater Houston, Texas metropolitan region, at nine (9) Continuous Ambient Monitoring Stations (CAMS) operated by the Texas Commission on Environmental Quality (TCEQ) between November 1, 2021 and October 31, 2022. The CAMS sites (CAMS 1, 8, 35, 45, 148, 403, 405, 410, and 1052) were chosen because their instrumentation includes measurements of PM2.5. These sites also provide continuous multi-parameter air-quality and meteorological measurements. Particulate matter (PM2.5, PM10) was sampled along with several trace gases, including ozone (O3), nitrogen oxides (NO, NO2, NOx), sulfur dioxide (SO2), and carbon monoxide (CO). The data set also contains standard surface meteorological parameters (temperature, humidity, pressure, wind speed, and wind direction). Several sites also include AutoGC-based measurements of volatile organic compounds (VOCs). Air monitoring instruments deployed at the selected sites comprise the following systems: BAM-1020 or TEOM (PM2.5), Thermo Scientific TEI 49i (O3), TEI 42i (NOx), and AutoGCs (VOCs). This data set is similar to the data included within the houairq5mX1.00 datastream, except for a few additional quality control steps. A systematic data cleaning and verification process was performed on the data set to ensure its quality and preparation for analysis. Removal of non-numeric status flags (e.g., [LIM], [QAS], [SPZ], [CAL], [PMA], [AQI], [SPN], [MAL]) was accomplished by employing rule-based string parsing to extract valid numerical values. Missing entries were set to -9999; however, invalid or anomalous values (e.g., 99999) were retained as originally reported by the TCEQ to preserve data provenance. The time sequence was verified for completeness, removal of duplicates, and uniformity at 5-minute intervals. Column labeling was standardized, and corresponding values were assessed for physical plausibility. All timestamps in the data set were reported in Coordinated Universal Time (UTC) as provided by the TCEQ. Further, the latitude and longitude coordinates were added for each CAMS site. A subset of the data (June 1–September 30, 2022) has been used in the following publication: Subba et al. 2025. “Implications of sea breeze circulations on boundary layer aerosols in the southern coastal Texas region.” EGUsphere 2025: 1–49, https://doi.org/10.5194/egusphere-2025-2659.

latitude↗

The Land Surface in Current and Planned MERRA Reanalysis Products

Current global atmospheric reanalysis products such as the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5), the NASA Modern-Era Retrospective analysis for Research and Applications version 2 (MERRA-2), and the Japanese Reanalysis for Three Quarters of a Century (JRA-3Q) provide estimates of land surface states and fluxes, including soil moisture, soil temperature, snow mass, latent and sensible heat fluxes, and runoff, that are widely used in research and applications. These land surface estimates are based on land surface process models and, depending on the reanalysis product, on precipitation observations or the assimilation of land surface observations of soil moisture, soil temperature, snow conditions, and screen-level air temperature and humidity from satellite observations and in situ measurements. In this presentation, we review the land surface modeling and data assimilation components of the suite of current and planned MERRA reanalysis products. In addition to MERRA-2, we will discuss the latest NASA reanalysis, MERRA for the 21st century (M21C), which is currently under production, as well as the development and planning of the next version of the MERRA reanalysis, tentatively labeled MERRA-3. In MERRA-2, observations-based precipitation data products are used to correct the precipitation falling on the land surface. Outside of the high-latitudes and Africa, the daily, 0.5-degree, gauge-based Climate Prediction Center (CPC) Unified (CPCU) product is used. In Africa, the pentad, 2.5-degree, satellite- and gauge-based CPC Merged Analysis of Precipitation (CMAP) product is used. Poleward of 62.5 degrees latitude, the land surface sees the precipitation generated by the atmospheric model in the cycling data assimilation system. This configuration provides improved soil moisture estimates compared to those of the original (version 1) MERRA estimates, which did not benefit from the use of precipitation observations. Moreover, the use of precipitation observations facilitates a seamless spin-up of the land surface initial conditions across the MERRA-2 production streams. The use of a gauge-only precipitation product in MERRA-2 across much of the globe, however, adversely impacts the quality of the MERRA-2 land surface estimates in regions with poor gauge coverage, including most of South America and Australia. Therefore, the forthcoming M21C reanalysis uses satellite- and gauge-based precipitation from the Integrated Multi-satellitE Retrievals for the Global Precipitation Measurement Mission (IMERG). This change results in significant improvements in the quality of the M21C soil moisture estimates in the Southern Hemisphere compared to those from MERRA-2. Planning for MERRA-3 focuses on the assimilation of soil moisture observations from the Soil Moisture Active Passive (SMAP) mission and the Advanced Scatterometer (ASCAT), along with snow cover area fraction observations from the Moderate Resolution Imaging Spectroradiometer (MODIS) to further improve the quality of the land surface estimates from the reanalysis. As a first step towards the assimilation of land surface observations in MERRA-3, the offline (land-only) M21C-Land reanalysis is currently under development as a supplemental M21C product that includes the assimilation of SMAP, ASCAT, and MODIS observations. Preliminary results from M21C and M21C-Land will be discussed in the context of MERRA-2 and plans for MERRA-3.

Rolf Reichle↗

Fluorescent Approaches to High Throughput Crystallography

We have shown that by covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and the presence of the probe at low concentrations does not affect the X-ray data quality or the crystallization behavior. The presence of the trace fluorescent label gives a number of advantages when used with high throughput crystallizations. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a dark background. Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Brightly fluorescent crystals are readily found against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. We are now testing the use of high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that kinetics leading to non-structured phases may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Preliminary experiments with test proteins have resulted in the extraction of a number of crystallization conditions from screening outcomes based solely on the presence of bright fluorescent regions. Subsequent experiments will test this approach using a wider range of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.↗