Search NASASearch

DOE OSTI · 2506686

Labeling sequential data from noisy annotations

Abstract

Crowdsourcing algorithms often work under the assumption that the data samples are independent. Recent work has shown that data dependence, such as temporal correlations in sequential data, can be leveraged to improve the label quality. Existing methods that exploit this special structure rely on third-order statistics of the annotator outputs to ensure the identifiability of key latent parameters, which are costly to acquire. This work proposes an approach for integrating crowdsourced annotations under the Dawid-Skene/Hidden Markov Model (DS-HMM) for sequential data based on second-order statistics, which naturally enjoys a lower sample complexity. An effective algorithm is proposed to tackle the challenging optimization problem associated with the proposed estimator. Numerical experiments showcase the effectiveness of the data labeling paradigm.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Marrinan, Timothy P., Ibrahim, Shahana, Fu, Xiao. 2024-08-26. Labeling sequential data from noisy annotations. https://doi.org/10.1109/sam60225.2024.10636383

Cite the original work for its findings. Save a collection to share your selection of sources.