Search NASA⌕ Search

SEARCH · Search NASA

Results for “label quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

RECON Label Quality Report

The final quality of any AI/ML system is directly related to the quality of the input data used to train the system. In this case, we are trying to build a reliable image classifier that can correctly identify electrical components in x-ray images. The classification confidence is directly related to the quality of the labels in the training data, which are used in developing the AI/ML classifier. Incorrect or incomplete labels can substantially hinder the performance of the system during the training process, as it tries to compensate for variations that should not exist. Image labels are entered by subject matter experts, and in general can be assumed to be correct. However, this is not a guarantee, so developing ways to measure label quality and help identify or reject bad labels is important, especially as the database continues to grow. Given the current size of the database, a full manual review of each component is not feasible. This report will highlight the current state of the “RECON” x-ray image database and summarize several recent developments to try to help ensure high quality labeling both now and in the future. Questions that we hope to answer with this development include: 1) Are there any components with incorrect labels? 2) Can we suggest labels for components that are marked “Unknown”? 3) What kind of overall confidence do we have in the quality of the existing labels? 4) What systems or procedures can we put in place to maximize label quality?

42 ENGINEERING↗

RectifHydPlus: Forty Year Hydropower Generation Reanalysis for Conterminous United States, Version 1.1.

This dataset contains monthly hydropower net-generation totals for 590 plants (each >10 MW) across the conterminous United States (CONUS) from 1980 to 2019. RectifHydPlus v1.1 includes one harmonized table of historical monthly generation—backfilled with observed monthly values where available—and two companion tables: (i) an estimates-only version with no backfill and (ii) a hydrological-control version that removes the effects of capacity and operational change. Each table comprises 23,600 records (590 plants × 40 years). The dataset was developed to address temporal gaps and inconsistencies in publicly available hydropower generation data as available through EIA-923 survey reports. Each record includes a quality label denoting the underlying proxy—from best (direct reservoir releases) to weakest (pattern copied from similar years). By combining the agency-reported survey records with observed and simulated hydrologic releases, RectifHydPlus offers complete, quality-labeled monthly estimates suitable for trend analysis and generation of hydropower generation inputs for energy-water modeling.

Turner, Sean [Oak Ridge National Laboratory (ORNL)↗

Domain Shift Analysis in Chest Radiographs Classification in a Veterans Healthcare Administration Population

This study aims to assess the impact of domain shift on chest X-ray classification accuracy and to analyze the influence of ground truth label quality and demographic factors such as age group, sex, and study year. We used a DenseNet121 model pre-trained MIMIC-CXR dataset for deep learning-based multi-label classification using ground truth labels from radiology reports extracted using the CheXpert and CheXbert Labeler. We compared the performance of the 14 chest X-ray labels on the MIMIC-CXR and Veterans Healthcare Administration chest X-ray dataset (VA-CXR). The validation of ground truth and the assessment of multi-label classification performance across various NLP extraction tools revealed that the VA-CXR dataset exhibited lower disagreement rates than the MIMIC-CXR datasets. Additionally, there were notable differences in AUC scores between models utilizing CheXpert and CheXbert. When evaluating multi-label classification performance across different datasets, minimal domain shift was observed in the unseen VA dataset, except for the label “Enlarged Cardiomediastinum.” The subgroup with the most significant variations in multi-label classification performance was study year. These findings underscore the importance of considering domain shift in chest X-ray classification tasks, paying particular attention to the temporality of the exam. Our study reveals the significant impact of domain shift and demographic factors on chest X-ray classification, emphasizing the need for improved transfer learning and robust model development. Addressing these challenges is crucial for advancing medical imaging research and improving patient care.

chest X-ray image classification↗

Labeling sequential data from noisy annotations

Crowdsourcing algorithms often work under the assumption that the data samples are independent. Recent work has shown that data dependence, such as temporal correlations in sequential data, can be leveraged to improve the label quality. Existing methods that exploit this special structure rely on third-order statistics of the annotator outputs to ensure the identifiability of key latent parameters, which are costly to acquire. This work proposes an approach for integrating crowdsourced annotations under the Dawid-Skene/Hidden Markov Model (DS-HMM) for sequential data based on second-order statistics, which naturally enjoys a lower sample complexity. An effective algorithm is proposed to tackle the challenging optimization problem associated with the proposed estimator. Numerical experiments showcase the effectiveness of the data labeling paradigm.

Marrinan, Timothy P.↗

A large expert-curated cryo-EM image dataset for machine learning protein particle picking

Cryo-electron microscopy (cryo-EM) is a powerful technique for determining the structures of biological macromolecular complexes. Picking single-protein particles from cryo-EM micrographs is a crucial step in reconstructing protein structures. However, the widely used template-based particle picking process is labor-intensive and time-consuming. Though machine learning and artificial intelligence (AI) based particle picking can potentially automate the process, its development is hindered by lack of large, high-quality labelled training data. To address this bottleneck, we present CryoPPP, a large, diverse, expert-curated cryo-EM image dataset for protein particle picking and analysis. It consists of labelled cryo-EM micrographs (images) of 34 representative protein datasets selected from the Electron Microscopy Public Image Archive (EMPIAR). The dataset is 2.6 terabytes and includes 9,893 high-resolution micrographs with labelled protein particle coordinates. The labelling process was rigorously validated through 2D particle class validation and 3D density map validation with the gold standard. The dataset is expected to greatly facilitate the development of both AI and classical methods for automated cryo-EM protein particle picking.

42 ENGINEERING↗

Introduction to Special Section: Machine Learning for Image-based Geologic Interpretation

Image-based geological interpretation has been a labor-intensive and time-consuming process because it requires well-trained geoscientists to identify geological structures, features, and textures from various types of images. These images include scanning electron microscopic images, optical microscopic images, optical photos, resistivity images, seismic volumes, remote-sensing images, etc. With fast-evolving machine learning (ML) technology and computing power in recent decades, computers can achieve nearhuman-level to super-human-level performance with scalable high efficiency in the computer vision field. These technological revolutions facilitated image-based geological interpretation in petroleum exploration and production. For example, a fault picking method applied to 3-D seismic volume data using deep learning can achieve superior performance in comparison to conventional auto-picking methods. In addition, under the new normal of low oil prices, the petroleum industry seeks cost-effective strategies such as automating traditionally labor-intensive processes. Nevertheless, the potential of applying ML to geological image interpretation is still facing a few key challenges including data scarcity, data distribution, poor data and/or label quality, data leakage, learning algorithms, model architecture, training methodologies, testing and evaluation metrics, hyper-parameters optimization, model drift, production deployment, and the like.

58 GEOSCIENCES↗

Physics-Based Method for Generating Fully Synthetic IV Curve Training Datasets for Machine Learning Classification of PV Failures

Classification machine learning models require high-quality labeled datasets for training. Among the most useful datasets for photovoltaic array fault detection and diagnosis are module or string current-voltage (IV) curves. Unfortunately, such datasets are rarely collected due to the cost of high fidelity monitoring, and the data that is available is generally not ideal, often consisting of unbalanced classes, noisy data due to environmental conditions, and few samples. In this paper, we propose an alternate approach that utilizes physics-based simulations of string-level IV curves as a fully synthetic training corpus that is independent of the test dataset. In our example, the training corpus consists of baseline (no fault), partial soiling, and cell crack system modes. The training corpus is used to train a 1D convolutional neural network (CNN) for failure classification. The approach is validated by comparing the model’s ability to classify failures detected on a real, measured IV curve testing corpus obtained from laboratory and field experiments. Results obtained using a fully synthetic training dataset achieve identical accuracy to those obtained with use of a measured training dataset. When evaluating the measured data’s test split, a 100% accuracy was found both when using simulations or measured data as the training corpus. When evaluating all of the measured data, a 96% accuracy was found when using a fully synthetic training dataset. The use of physics-based modeling results as a training corpus for failure detection and classification has many advantages for implementation as each PV system is configured differently, and it would be nearly impossible to train using labeled measured data.

Hopwood, Michael W. (ORCID:0000000161901767)↗

Improving Building Footprint Extraction Using NAIP and 3DEP Lidar Derived Features with Deep Learning

Accurate building footprint extraction is critical for applications ranging from population estimation to disaster management. Although optical imagery provides detailed spectral information, it often struggles with shadows, occlusions, and background clutter in dense urban environments. Lidar data, by contrast, offer precise elevation and structural attributes but face challenges such as variable point density and noise. This study integrates multispectral imagery from the U.S. Department of Agriculture (USDA) National Agriculture Imagery Program (NAIP) with lidar-derived feature height and intensity from the U.S. Geological Survey (USGS) 3D Elevation Program (3DEP) to improve footprint extraction using a U-Net–based deep learning model. A six-band input stack (RGB, near-infrared, height, intensity) was developed, normalized, and tiled for training and evaluation against Microsoft Global Building Footprints (GBF). Results from the Houston, TX test site show that the six-band model achieved a precision of 0.86, recall of 0.88, F1 score of 0.87, and Intersection-over-Union (IoU) of 0.76, consistently outperforming four-band baselines by reducing false positives while maintaining sensitivity. Predictions on withheld Houston tiles confirmed strong within-region generalization, yielded a precision of 0.78, recall of 0.81, F1 score of 0.79, and IoU of 0.66. Qualitative analysis further revealed limitations stemming from both training label quality and vegetation–building confusion. These findings demonstrate the complementary value of integrating spectral and structural information for robust building footprint extraction and how domain adaptation strategies can be used to enhance cross-regional transferability.

Liu, Jung Kuan [United States Geological Survey (U↗

Learning instrument invariant characteristics for generating high-resolution global coral reef maps

Coral reefs are one of the most biologically complex and diverse ecosystems within the shallow marine environment. Unfortunately, these underwater ecosystems are threatened by a number of anthropogenic challenges, including ocean acidification and warming, overfishing, and the continued increase of marine debris in oceans. This requires a comprehensive assessment of the world's coastal environments, including a quantitative analysis on the health and extent of coral reefs and other associated marine species, as a vital Earth Science measurement. However, limitations in observational and technological capabilities inhibit global sustained imaging of the marine environment. Harmonizing multimodal data sets acquired using different remote sensing instruments presents additional challenges, thereby limiting the availability of good quality labeled data for analysis. In this work, we develop a deep learning model for extracting domain invariant features from multimodal remote sensing imagery and creating high-resolution global maps of coral reefs by combining various sources of imagery and limited hand-labeled data available for certain regions. This framework allows us to generate, for the first time, coral reef segmentation maps at 2-meter resolution, which is a significant improvement over the kilometer-scale state-of-the-art maps. Additionally, this framework doubles accuracy and IoU metrics over baselines that do not account for domain invariance.

Domain Adaptation↗

FloodPlanet: High-Resolution Commercial Imagery for Training and Validation of Deep Learning-Based Models of Inundation Extent

Flooding events are becoming increasingly frequent worldwide and are known to cause extensive damage. Public optical and radar satellite imagery can be used to detect large areas of inundation in rural areas, however, long revisit times and coarse spatial resolution limit applications for short-lived events and urban areas. Commercial constellations such as those operated by Planet offer increased spatial and temporal resolution and can supplement mapping efforts to provide more information to disaster response, relief, and mitigation efforts. Deep learning requires high quality labeled data for training across coincident sensors. The FloodPlanet dataset presented here contains labeled surface water for 18 events across the world based on Planetscope imagery with coincident Harmonized Landsat Sentinel-2 ( HLS) or Sentinel-1 and builds upon the previously existing Sen1Floods11, xBD, and NASA Sentinel-1 datasets. Sen1Floods11 includes 4,831 512x512 pixel overlapping tiles of coincident Sentinel-1 and Sentinel-2 data observing 11 flood events across the world from 2017-2019. The dataset contains a combination of automated and hand-labeled surface water for use in training and validation of inundation modeling efforts. The xBD dataset identifies flood-damaged buildings and indicates the scale of damage to each (none, minor, moderate, and major) from four flood events which occurred in the United States, India, Nepal, and Bangladesh from the same time period. The NASA dataset contains hand-labeled water bodies observed in Sentinel-1 imagery during five flood events within the 2017-2019 period. The effort presented here utilizes observations from these previously investigated flood events to generate labels of surface water at the 3-5m spatial resolution provided by Planetscope and facilitate the comparison between public and commercial data. A data pipeline was built which uses clustering algorithms to pick the most suitable overlapping chips between the public data and PlanetScope data for manual labeling. Labels were created manually using NASA’s ImageLabeler tool and include areas of high- and low-confidence water. The high confidence designation is reserved for areas of open, unobstructed water while low confidence is used for areas of suspected water beneath vegetation, clouds, or cloud shadows. Expected to be released in late 2022, the FloodPlanet dataset will include tiled imagery with a unique ID for each 1024x1024 pixel tile, 7 bands of HLS data, and high- and low-confidence flood labels in both shapefile and tiff formats. The authors will follow Spatial Temporal Access Catalog (STAC) guidelines to release FloodPlanet on the Radiant Earth ML hub, which hosts public datasets for machine learning.

Alexander Melancon↗

Scatter-Reducing Sounding Filtration Using a Genetic Algorithm and Mean Monthly Standard Deviation

Retrieval algorithms like that used by the Orbiting Carbon Observatory (OCO)-2 mission generate massive quantities of data of varying quality and reliability. A computationally efficient, simple method of labeling problematic datapoints or predicting soundings that will fail is required for basic operation, given that only 6% of the retrieved data may be operationally processed. This method automatically obtains a filter designed to reduce scatter based on a small number of input features. Most machine-learning filter construction algorithms attempt to predict error in the CO2 value. By using a surrogate goal of Mean Monthly STDEV, the goal is to reduce the retrieved CO2 scatter rather than solving the harder problem of reducing CO2 error. This lends itself to improved interpretability and performance. This software reduces the scatter of retrieved CO2 values globally based on a minimum number of input features. It can be used as a prefilter to reduce the number of soundings requested, or as a post-filter to label data quality. The use of the MMS (Mean Monthly Standard deviation) provides a much cleaner, clearer filter than the standard ABS(CO2-truth) metrics previously employed by competitor methods. The software's main strength lies in a clearer (i.e., fewer features required) filter that more efficiently reduces scatter in retrieved CO2 rather than focusing on the more complex (and easily removed) bias issues.

Mandrake, Lukas↗

Hierarchical Convolutional Neural Networks for Event Classification on PMU Measurements

Event classification is one of the central components of automated disturbance analysis based on PMU measurements. Obtaining high-quality event labels remains a challenge for supervised learning-based classification of local and system-wide events in power grids due to its labor-intensive requirement. We present a sensitivity study considering rapidly refined, partially and fully inspected event labels that leads to evidence that hierarchical convolutional neural networks (HCNNs) outperform traditional classification models regardless of the quality of the available event labels. Furthermore, it is demonstrated that performance similar to the one obtained using entirely domain-driven labeling can be achieved as long as the involved expert does not mislabel more than ~5% of the event data captured by PMU measurements.

47 OTHER INSTRUMENTATION↗

A data-centric weak supervised learning for highway traffic incident detection

Using the data from loop detector sensors for near-real-time detection of traffic incidents on highways is crucial to averting major traffic congestion. While recent supervised machine learning methods offer solutions to incident detection by leveraging human-labeled incident data, the false alarm rate is often too high to be used in practice. Specifically, the inconsistency in the human labeling of the incidents significantly affects the performance of supervised learning models. To that end, we focus on a data-centric approach to improve the accuracy and reduce the false alarm rate of traffic incident detection on highways. We develop a weak supervised learning workflow to generate high-quality training labels for the incident data without the ground truth labels, and we use those generated labels in the supervised learning setup for final detection. This approach comprises three stages. First, we introduce a data preprocessing and curation pipeline that processes traffic sensor data to generate high-quality training data through leveraging labeling functions, which can be domain knowledge-related or simple heuristic rules. Second, we evaluate the training data generated by weak supervision using three supervised learning models-random forest, k-nearest neighbors, and a support vector machine ensemble-and long short-term memory classifiers. The results show that the accuracy of all of the models improves significantly after using the training data generated by weak supervision. Third, we develop an online real-time incident detection approach that leverages the model ensemble and the uncertainty quantification while detecting incidents. Finally, we show that our proposed weak supervised learning workflow achieves a high incident detection rate (0.90) and low false alarm rate (0.08).

97 MATHEMATICS AND COMPUTING↗

A Data-Driven Framework for Power System Event Type Identification via Safe Semi-Supervised Techniques

Herein this paper investigates the use of phasor measurement unit (PMU) data with deep learning techniques to construct real-time event identification models for transmission networks. Increasing penetration of distributed energy resources represents a great opportunity to achieve decarbonization, as well as challenges in systematic situational awareness. When high-resolution PMU data and sufficient manually recorded event labels are available, the power event identification problem is defined as a statistical classification problem that can be solved by numerous cutting-edge classifiers. However, in real grids, collecting tremendous high-quality event labels is quite expensive. Utilities frequently have a large number of event records without in-depth details (i.e., unlabeled events). To bridge this gap, we propose a novel semi-supervised learning-based method to improve the performance of event classifiers trained with a limited number of labeled events by exploiting the information from massive unlabeled events. In other words, compared to existing data-driven methods, our method requires only a small portion of labeled data to achieve a similar level of accuracy. Meanwhile, this work discusses and addresses the performance degradation caused by class distribution mismatch between the training set and the real applications. Based on the proposed safe learning mechanism, our model does not directly use all unlabeled events during model training, but selectively uses them through a comprehensive evaluation procedure. Numerical studies on a sizable PMU dataset have been used to validate the performance of the proposed method.

42 ENGINEERING↗

High-Dimensional Bayesian Optimization via Semi-Supervised Learning with Optimized Unlabeled Data Sampling

We introduce a novel semi-supervised learning approach, named Teacher-Student Bayesian Optimization (TSBO ), integrating the teacher-student paradigm into BO to minimize expensive labeled data queries for the first time. TSBO incorporates a teacher model, an unlabeled data sampler, and a student model. The student is trained on unlabeled data locations generated by the sampler, with pseudo labels predicted by the teacher. The interplay between these three components implements a unique selective regularization to the teacher in the form of student feedback. This scheme enables the teacher to predict high-quality pseudo labels, enhancing the generalization of the GP surrogate model in the search space. To fully exploit TSBO , we propose two optimized unlabeled data samplers to construct effective student feedback that well aligns with the objective of Bayesian optimization. Furthermore, we quantify and leverage the uncertainty of the teacher-student model for the provision of reliable feedback to the teacher in the presence of risky pseudo-label predictions. TSBO demonstrates significantly improved sample-efficiency in several global optimization tasks under tight labeled data budgets. The implementation is available at https://github.com/reminiscenty/TSBO-Official.

Yin, Yuxuan↗

Online Power System Event Detection via Bidirectional Generative Adversarial Networks

Accurate and speedy detection of power system events is critical to enhancing the reliability and resiliency of power systems. Although supervised deep learning algorithms show great promise in identifying power system events, they require a large volume of high-quality event labels for training. This paper develops a bidirectional anomaly generative adversarial network (GAN)-based algorithm to detect power system events using streaming PMU data, which does not rely on a huge amount of event labels. By introducing conditional entropy constraint in the objective function of GAN and graph signal processing-based PMU sorting technique, our proposed algorithm significantly outperforms state-of-the-art event detection algorithms in terms of accuracy. To facilitate the adoption of the proposed algorithm, a prototype online platform is also developed using Apache Hadoop, Kafka, and Spark to enable real-time event detection. Here, the accuracy and computational efficiency of the proposed algorithm are validated using a large-scale real-world PMU dataset from the Eastern Interconnection of the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Dynamic Reconfiguration of Switchgrass Proteomes in Response to Rust ( Puccinia novopanici ) Infection

Switchgrass (Panicum virgatum L.) can be infected by the rust pathogen (Puccinia novopanici) and results in lowering biomass yields and quality. Label-free quantitative proteomics was conducted on leaf extracts harvested from non-infected and infected plants from a susceptible cultivar (Summer) at 7, 11, and 18 days after inoculation (DAI) to follow the progression of disease and evaluate any plant compensatory mechanisms to infection. Some pustules were evident at 7 DAI, and their numbers increased with time. However, fungal DNA loads did not appreciably change over the course of this experiment in the infected plants. In total, 3830 proteins were identified at 1% false discovery rate, with 3632 mapped to the switchgrass proteome and 198 proteins mapped to different Puccinia proteomes. Across all comparisons, 1825 differentially accumulated switchgrass proteins were identified and subjected to a STRING analysis using Arabidopsis (A. thaliana L.) orthologs to deduce switchgrass cellular pathways impacted by rust infection. Proteins associated with plastid functions and primary metabolism were diminished in infected Summer plants at all harvest dates, whereas proteins associated with immunity, chaperone functions, and phenylpropanoid biosynthesis were significantly enriched. At 18 DAI, 1105 and 151 proteins were significantly enriched or diminished, respectively. Many of the enriched proteins were associated with mitigation of cellular stress and defense.

59 BASIC BIOLOGICAL SCIENCES↗