Search NASA⌕ Search

DOE OSTI · 1827039

The value of human data annotation for machine learning based anomaly detection in environmental systems

Abstract

Anomaly detection is the process of identifying unexpected data samples in datasets. Automated anomaly detection is either performed using supervised machine learning models, which require a labelled dataset for their calibration, or unsupervised models, which do not require labels. While academic research has produced a vast array of tools and machine learning models for automated anomaly detection, the research community focused on environmental systems still lacks a comparative analysis that is simultaneously comprehensive, objective, and systematic. This knowledge gap is addressed for the first time in this study, where 15 different supervised and unsupervised anomaly detection models are evaluated on 5 different environmental datasets from engineered and natural aquatic systems. To this end, anomaly detection performance, labelling efforts, as well as the impact of model and algorithm tuning are taken into account. As a result, our analysis reveals the relative strengths and weaknesses of the different approaches in an objective manner without bias for any particular paradigm in machine learning. Most importantly, our results show that expert-based data annotation is extremely valuable for anomaly detection based on machine learning.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Russo, Stefania, Besmer, Michael D., Blumensaat, Frank, Bouffard, Damien, Disch, Andy, Hammes, Frederik, Hess, Angelika, Lürig, Moritz, Matthews, Blake, Minaudo, Camille, Morgenroth, Eberhard, Tran-Khac, Viet, Villez, Kris. 2021-09-27. The value of human data annotation for machine learning based anomaly detection in environmental systems. https://doi.org/10.1016/j.watres.2021.117695

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Water4Energy Step-1 Band-M Ready-to-Train Samples for TVA Weeks-to-Years Prediction, Version 0

AI-ready Band-M (monthly) labelled training pack for the Water4Energy Genesis Task-1 project on weeks-to-years prediction of Tennessee Valley temperature and precipitation. The deposit includes leakage-aware issue-time samples (samples_M_v0.nc; N=486), train-only scalers, issue-time split table, supporting monthly panels, and Python generation scripts to recreate the pack from the companion Tier-1 raw observation collection (https://doi.org/10.13139/ORNLNCCS/3398576). Each sample pairs a 12-month lookback of teleconnection indices and SST box anomalies with TVA-mean ERA5 anomaly targets (t2m, tp, msl) at leads 1–3 months.

54 ENVIRONMENTAL SCIENCES↗

Multi-Angle Snowflake Camera, particle analysis

The c1 level data product for the Mutli-Angle Snowflake Camera contains snowflake fall speeds and particle size, among other analysis for images associated with each hydrometeor.

54 ENVIRONMENTAL SCIENCES↗

Multi-Angle Snowflake Camera, time bins

The c1 level data product for the Mutli-Angle Snowflake Camera contains snowflake fall speeds and particle size, among other analysis for images associated with each hydrometeor.

54 ENVIRONMENTAL SCIENCES↗