Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset Annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Publicly Available, Annotated Dataset for Naturalistic Driving Study and Computer Vision Algorithm Development

Oak Ridge National Laboratory developed and implemented a data collection effort to create a dataset for use in evaluating and testing algorithms for analyzing driver behavior under controlled settings for support of the Federal Highway Administration’s Exploratory Advanced Research Program. This collection is called the ORNL Naturalistic Driving Study Sample (ONDSS). The dataset is designed to emulate aspects of the Second Strategic Highway Research Project (SHRP2), which contained a massive naturalistic driving study (NDS) with over 3000 drivers between 2010 and 2013 using their personal vehicles, with over 4300 person-years of data collected [HANKEY].

42 ENGINEERING↗

Multi-Species Complex and Standard Metabolomic Samples with Verified Truth Annotations Dataset

This dataset contains 4523251 (~6.35 GB) metabolite-spectra matches following identification with CoreMS. Data were manually curated as true positives, true negatives, or unknowns. Calculations for spectral similarity scores were carried out with two methods for a total of ~12.7 GB (6.35 * 2) of data. They are all .tsv files, though can easily be changed to .txt. The file types are: * human cerebrospinal fluid (CSF), human blood plasma human urine: already published here https://www.nature.com/articles/s41597-021-00894-y, • purchased FAMES standards • fungi species (A. niger, A. nidulans, T. reesei) • soil crust

59 BASIC BIOLOGICAL SCIENCES↗

Acoustic Rocket Signatures Collected by Smartphones

Rockets generate complex acoustic signatures that can be detected over a thousand kilometers from their source. While many far-field acoustic rocket signatures have been collected and released to the public, very few signatures collected at distances less than 100 km are available. This work presents a curated and annotated dataset of acoustic signatures of 243 rocket launches collected by a network of smartphones stationed at distances between 10 and 70 km from the launch sites, resulting in 1089 individual recordings. Due to the frequency dependence of atmospheric attenuation and the relatively short propagation distances, higher-frequency features not preserved in most publicly available data are observed. The signals are time-aligned to allow for different segments of the signal (ignition, launch, trajectory, chronology) to be more easily examined and compared. Initial analysis of the features of these rocket launch stages is performed, observed features are compared to those found in the existing literature, and comparisons between signals from launches of different rocket types are made. The dataset is annotated and made available to the public to aid future analysis of the characteristics and source mechanisms of rocket acoustics as well as applications such as rocket detection and classification models.

33 ADVANCED PROPULSION SYSTEMS↗

Improving microstructures segmentation via pretraining with synthetic data

Image analysis of material microstructures through microscopy is an integral capability in the field of materials science. The topological and chemical information obtained through microscopy allow us to draw vital connections between material microstructures, properties, and processing. While scanning electron microscopy (SEM) is able to yield a considerable wealth of information interpretable by the intuition of experts, there has been considerable interest in using machine learning, convolutional neural networks (CNNs) in particular, for such image analysis task. Training CNNs for an image analysis task requires a large annotated dataset. However, in many materials science applications, obtaining a large annotated dataset is cost and labor intensive. In this work, we study the use of synthetic data to enlarge the available annotated experimental data of uranium oxide. We utilize a modified Potts model to simulate uranium oxide particles with morphologies similar to those observed experimentally. We then leverage an image-to-image translation model to synthesize the simulated particles as if they are acquired with SEM. Through this process, we obtain pairs of particle images and their corresponding SEM representations, which corresponds to pairs of annotations and images. Unlike previous works, we leverage synthetic data for pretraining a CNN model prior, and finetune that model further with experimental data. We experimentally demonstrate that using synthetic data as incremental learning process benefits the overall performance compared to training a model on combined synthetic and experimental data.

36 MATERIALS SCIENCE↗

SFA-VirOmics

PNNL's Soil Microbiome Science Focus Area (SFA) is focused on understanding the basic biology underpinning how interactions among various soil microbial community members, across trophic levels, lead to the emergence of community functions. Moisture, in particular, drives microbial interactions and influences everything from cell function to substrate fate within soils. The group predicts this results in repeatable, predictable phenotypes. The sum of these phenotypes comprises the “soil metaphenome”. Understanding how the soil metaphenome shifts in response to moisture will provide a basis for modeling and predicting these shifts in reaction network responses. Visit the PNNL Soil Microbiome Science Focus Area Program homepage for more information. PNNL’s Soil Microbiome SFA Virome Dataset Annotation page is an extension to the PNNL Soil Microbiome SFA repository on DataHub allows for exploring and downloading integrated experimental omics dataset annotations, associated experimental metadata, pre- and post-processed viral data files, and other associated materials directly related to experimental viral project data.

59 BASIC BIOLOGICAL SCIENCES↗

Automated Global-Scale Detection and Characterization of Anthropogenic Activity using Multi-Source Satellite-Based Remote Sensing Imagery

Satellite-based remote sensing imagery is an effective means for detecting objects and structures in support of many applications. However, detecting the spatial and temporal bounds of a specific activity in satellite imagery is inherently more complex and research in this area is nascent. One reason for this is that describing an activity implies defining both spatial and temporal bounds and while activity is inherently continuous in nature, the geospatial (imagery) time series for any particular swath of ground provided by satellite imagery is relatively sparse and discrete in comparison. The IARPA Space-Based Machine Automated Recognition Technique (SMART)1 program is the first large-scale research program to target advancing the state of the art for automatically detecting, characterizing, and monitoring large-scale anthropogenic activity in global, multispectral satellite imagery. The program has two primary research objectives: 1) the “harmonization” of multiple imagery sources and 2) automated reasoning at scale to detect, characterize, and monitor activities of interest. This paper provides details on the goals, dataset, metrics, and lessons learned of the IARPA SMART program. By releasing the annotated dataset, the program aims to foster additional research in this area by the community at large.

Hirsh R Goldberg↗

CHQ- SocioEmo: Identifying Social and Emotional Support Needs in Consumer-Health Questions

General public, often called consumers, are increasingly seeking health information online. To be satisfactory, answers to health-related questions often have to go beyond informational needs. Automated approaches to consumer health question answering should be able to recognize the need for social and emotional support. Recently, large scale datasets have addressed the issue of medical question answering and highlighted the challenges associated with question classification from the standpoint of informational needs. However, there is a lack of annotated datasets for the non-informational needs. We introduce a new dataset for non-informational support needs, called CHQ-SocioEmo. The Dataset of Consumer Health Questions was collected from a community question answering forum and annotated with basic emotions and social support needs. This is the first publicly available resource for understanding non-informational support needs in consumer health-related questions online. We benchmark the corpus against multiple state-of-the-art classification models to demonstrate the dataset’s effectiveness.

60 APPLIED LIFE SCIENCES↗

Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object Detection

Object detection in remote sensing demands extensive, high-quality annotations—a process that is both labor-intensive and time-consuming. In this work, we introduce a real-time active learning and semi-automated labeling framework that leverages foundation models to streamline dataset annotation for object detection in remote sensing imagery. For example, by integrating a Segment Anything Model (SAM), our approach generates mask-based bounding boxes that serve as the basis for dual sampling: (a) uncertainty estimation to pinpoint challenging samples, and (b) diversity assessment to ensure broad data coverage. Furthermore, our Dynamic Box Switching Module (DBS) addresses the well-known cold start problem for object detection models by replacing its suboptimal initial predictions with SAM-derived masks, thereby enhancing early-stage localization accuracy. Extensive evaluations on multiple remote sensing datasets plus a real-world user study, demonstrate that our framework not only reduces annotation effort, but also significantly boosts detection performance compared to traditional active learning sampling methods. The code for training and the user interface will be made available.

Burges, Marvin [ORNL] (ORCID:0000000312690769)↗

MaizeMine: A Data Mining Warehouse for the Maize Genetics and Genomics Database

MaizeMine is the data mining resource of the Maize Genetics and Genome Database (MaizeGDB; http://maizemine.maizegdb.org). It enables researchers to create and export customized annotation datasets that can be merged with their own research data for use in downstream analyses. MaizeMine uses the InterMine data warehousing system to integrate genomic sequences and gene annotations from the Zea mays B73 RefGen_v3 and B73 RefGen_v4 genome assemblies, Gene Ontology annotations, single nucleotide polymorphisms, protein annotations, homologs, pathways, and precomputed gene expression levels based on RNA-seq data from the Z. mays B73 Gene Expression Atlas. MaizeMine also provides database cross references between genes of alternative gene sets from Gramene and NCBI RefSeq. MaizeMine includes several search tools, including a keyword search, built-in template queries with intuitive search menus, and a QueryBuilder tool for creating custom queries. The Genomic Regions search tool executes queries based on lists of genome coordinates, and supports both the B73 RefGen_v3 and B73 RefGen_v4 assemblies. The List tool allows you to upload identifiers to create custom lists, perform set operations such as unions and intersections, and execute template queries with lists. When used with gene identifiers, the List tool automatically provides gene set enrichment for Gene Ontology (GO) and pathways, with a choice of statistical parameters and background gene sets. With the ability to save query outputs as lists that can be input to new queries, MaizeMine provides limitless possibilities for data integration and meta-analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Semi-automatic image annotation using 3D LiDAR projections and depth camera data

Efficient image annotation is necessary to utilize deep learning object recognition neural networks in nuclear safeguards, such as for the detection and localization of target objects like nuclear material containers (NMCs). This capability can help automate the inventory accounting of different types of NMCs within nuclear storage facilities. The conventional manual annotation process is labor-intensive and time-consuming, hindering the rapid deployment of deep learning models for NMC identifications. This paper introduces a novel semi-automatic method for annotating 2D images of nuclear material containers (NMCs) by combining 3D light detection and ranging (LiDAR) data with color and depth camera images collected from a handheld scan system. The annotation pipeline involves an operator manually marking new target objects on a LiDAR-generated map, and projecting these 3D locations to images, thereby automatically creating annotations from the projections. The semi-automatic approach significantly reduces manual efforts and the expertise in image annotation that is required to perform the task, allowing deep learning models to be trained on-site within a few hours. The paper compares the performance of models trained on datasets annotated through various methods, including semi-automatic, manual, and commercial annotation services. The evaluation demonstrates that the semi-automatic annotation method achieves comparable or superior results, with a mean average precision (mAP) above 0.9, showcasing its efficiency in training object recognition models. Additionally, the paper explores the application of the proposed method to instance segmentation, achieving promising results in detecting multiple types of NMCs in various formations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

A Natural Language Understanding Approach for Digitizing Aircraft Ground Taxi Instructions

Advancements in natural language processing (NLP) technologies offer a unique opportunity to furnish aircraft crews, primarily pilots, with digital instructions for taxiing operations. Digital taxi instructions, delivered either as text or graphics, can streamline taxiing procedures, thereby reducing radio congestion, minimizing communication errors, and enhancing aircraft monitoring. Techniques used for natural language understanding (NLU), a subset of NLP focused on machine comprehension of natural language, can extract taxi instructions directly from verbal radio communications. This capability paves the way for implementing a digital taxi communication framework with minimal adjustments to the existing air traffic controller operations. This paper delves into a novel application of NLU: the automated generation of digital taxi instructions from air traffic controller speech. We detail the development of an annotation scheme to represent aircraft ground traffic communications within the US National Airspace System (NAS), employing intent classification (IC) and slot filling (SF) to extract taxi instructions using NLU models. Several neural network models were trained on a dataset annotated with our scheme, achieving notable accuracy and F1 scores. Our research demonstrates the feasibility of using NLU to automatically generate digital taxi instructions, showcasing its potential to streamline the implementation of digital taxi communications.

LSTM↗

A Natural Language Understanding Approach for Digitizing Aircraft Ground Taxi Instructions

Advancements in natural language processing (NLP) technologies offer a unique opportunity to furnish aircraft crews, primarily pilots, with digital instructions for taxiing operations. Digital taxi instructions, delivered either as text or graphics, can streamline taxiing procedures, thereby reducing radio congestion, minimizing communication errors, and enhancing aircraft monitoring. Techniques used for natural language understanding (NLU), a subset of NLP focused on machine comprehension of natural language, can extract taxi instructions directly from verbal radio communications. This capability paves the way for implementing a digital taxi communication framework with minimal adjustments to the existing air traffic controller operations. This paper delves into a novel application of NLU: the automated generation of digital taxi instructions from air traffic controller speech. We detail the development of an annotation scheme to represent aircraft ground traffic communications within the US National Airspace System (NAS), employing intent classification (IC) and slot filling (SF) to extract taxi instructions using NLU models. Several neural network models were trained on a dataset annotated with our scheme, achieving notable accuracy and 𝐹1 scores. Our research demonstrates the feasibility of using NLU to automatically generate digital taxi instructions, showcasing its potential to streamline the implementation of digital taxi communications.

ATC↗

Extracting Material Property Measurement Data from Scientific Articles

Machine learning-based prediction of material properties is often hampered by the lack of sufficiently large training datasets. The majority of such measurement data is embedded in scientific literature and the ability to automatically extract these data is essential to support the development of reliable property prediction methods. In this work, we describe a methodology for an automatic property extraction framework using material solubility as the target property. We create an annotated dataset containing tags for solubility-related entities using a combination of regular expressions and manual tagging. We then compare five entity recognition models leveraging both token-level and span-level architectures on the task of classifying solute names, solubility values, and solubility units. Additionally, we explore a novel pretraining approach that leverages automated chemical name and quantity extraction tools to generate large datasets that do not rely on intensive manual effort. Finally, we perform an analysis to identify the causes of classification errors.

Panapitiya, Gihan U.↗

Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data

Zero-shot and prompt-based models have excelled at visual reasoning tasks by leveraging large-scale natural image corpora, but they often fail on sparse and domain-specific scientific image data. We introduce Zenesis, a no-code interactive computer vision platform designed to reduce data readiness bottlenecks in scientific imaging workflows. Zenesis integrates lightweight multimodal adaptation for zero-shot inference on raw scientific data, human-in-the-loop refinement, and heuristic-based temporal enhancement. We validate our approach on Focused Ion Beam Scanning Electron Microscopy (FIB-SEM) datasets of catalyst-loaded membranes. Zenesis outperforms baselines, achieving an average accuracy of 0.947, Intersection over Union (IoU) of 0.858, and Dice score of 0.923 on amorphous catalyst samples; and 0.987 accuracy, 0.857 IoU, and 0.923 Dice on crystalline samples. These results represent a significant performance gain over conventional methods such as Otsu thresholding and standalone models like the Segment Anything Model (SAM). Zenesis enables effective image segmentation in domains where annotated datasets are limited, offering a scalable solution for scientific discovery.

Mukherjee, Shubhabrata↗

Patch-Based Convolutional Neural Networks for Multiple Microstructural Features Detection in FIB-SEM Micrographs of Irradiated Nuclear Fuel

Focused ion beam scanning electron microscopy (FIB-SEM) tomography has increasingly been utilized for acquiring three-dimensional (3D) microstructure features at the sub-micron scale in irradiated nuclear materials. This technique involves sequential ion beam slicing followed by electron beam imaging and compositional mapping using energy dispersive spectroscopy (EDS). Despite its growing use, several challenges persist. These include the time-intensive nature of data collection of EDS data, difficulties in distinguishing between various microstructures, and issues with image alignment. These challenges currently limit the broader application of FIB-SEM tomography in the field. To overcome these limitations, we propose using convolutional neural networks (CNNs) to automate microstructure identification in SEM images. Our study introduces a new framework for identifying microstructures in irradiated U-10Zr (wt. %) metallic fuel with limited annotated data. The framework includes the creation of a reliable annotated dataset with paired SEM and ground truth data from EDS maps, the applications of CNNs for microstructure identification, and the validation of model performance. Specifically, we employed the Segment Anything Model (SAM) to align SEM images with corresponding EDS maps and focused ion beam (FIB) tomography SEM data. We evaluate several models, including Patch-based U-Net, Attention U-Net, and Residual U-Net, finding that patch-based U-Net exhibits superior segmentation performance and consistency. This approach reduces reliance on EDS detectors and aids in accelerating nuclear material analysis process, highlighting the potential of advanced deep learning techniques to improve microstructural understanding in nuclear material. This is the first framework to integrate SAM and Patch-based CNN models for semantic segmentation of irradiated nuclear materials, with potential applicability to other tomography datasets.

36 - MATERIALS SCIENCE↗

Automatic Crack Segmentation and Feature Extraction in Electroluminescence Images of Solar Modules

The effect of cracks in solar cells on the long-term degradation of photovoltaic (PV) modules remains to be determined. To investigate this effect in future studies, it is necessary to quantitatively describe the crack features (e.g., length) and correlate them with module power loss. Electroluminescence (EL) imaging is a common technique for identifying cracks. However, it is currently challenging and time-consuming to identify cracks in a large number of EL images and quantify complex crack features by human inspection. This article introduces a fast semantic segmentation method (~0.18 s/cell) to automatically segment cracks from EL images and algorithms to extract crack features. Here we fine-tuned a UNet neural network model using pretrained VGG16 as the encoder and obtained an average F1 score of 0.875 and an intersection over union score of 0.782 on the testing set. With cracks and busbars segmented, we developed algorithms for extracting crack features, including the crack-isolated area, the brightness inside the isolated area, and the crack length. We also developed an automatic preprocessing tool for cropping individual cell images from EL images of PV modules (~0.72 s/module). Our codes are published as open-source an software, and our annotated dataset composed of various types of cells is published as a benchmark for crack segmentation in EL images.

14 SOLAR ENERGY↗

SpaceNet 9—Cross-Sensor Alignment of Optical and SAR Imagery

Precise registration of high-resolution synthetic aperture radar (SAR) and optical imagery is necessary for realizing the full potential and benefits of multimodal image analysis. However, two significant challenges presently exist. First, there is a lack of annotated datasets and benchmarks available for high-resolution SAR–optical image registration. Second, an assessment of efficient and reliable image registration methods that can precisely align these modalities is lacking. Here, we present a holistic description of the SpaceNet 9 Challenge and its results. We present a description of the dataset and baseline algorithm along with the results of the challenge, including a description of the winning algorithms. We release the SpaceNet 9 dataset along with open-sourcing the winning algorithms and baseline. The objective of SpaceNet 9 was to compute a dense displacement map that indicates the shift needed to align pixels in an optical image to the pixels in a SAR image. The challenge launched in April 2025 and was active for approximately two months. The top five solutions reduced image alignment error from approximately 34 m to under 13 m for public and private test data, with the best results obtaining a registration error of only 8.5 and 6.7 m on the public testing and private testing dataset, respectively. Usage of pretrained image matching models, robust outlier rejection with RANSAC, and estimating local displacement were common among the top solutions. The results of this challenge provide insight into high-resolution SAR–optical image registration and offer opportunities for future benchmarking in this domain. The baseline algorithm, winning solutions, and datasets are available at https://spacenet.ai/sn9-challenge/.

benchmark datasets↗

Diagnosing Malaria Patients with Plasmodium falciparum and vivax Using Deep Learning for Thick Smear Images

We propose a new framework, PlasmodiumVF-Net, to analyze thick smear microscopy images for a malaria diagnosis on both image and patient-level. Our framework detects whether a patient is infected, and in case of a malarial infection, reports whether the patient is infected by Plasmodium falciparum or Plasmodium vivax. PlasmodiumVF-Net first detects candidates for Plasmodium parasites using a Mask Regional-Convolutional Neural Network (Mask R-CNN), filters out false positives using a ResNet50 classifier, and then follows a new approach to recognize parasite species based on a score obtained from the number of detected patches and their aggregated probabilities for all of the patient images. Reporting a patient-level decision is highly challenging, and therefore reported less often in the literature, due to the small size of detected parasites, the similarity to staining artifacts, the similarity of species in different development stages, and illumination or color variations on patient-level. We use a manually annotated dataset consisting of 350 patients, with about 6000 images, which we make publicly available together with this manuscript. Our framework achieves an overall accuracy above 90% on image and patient-level.

60 APPLIED LIFE SCIENCES↗