Search NASA⌕ Search

DOE OSTI · 1894779

Evaluating Deception Detection Model Robustness To Linguistic Variation

Abstract

With the increasing use of automated, machine learning-driven tools and the downstream impact that algorithmic judgements can have, it is critical to develop models that are robust to evolving or manipulated inputs. Evaluating the reliability of multimodal models across linguistic variations to understand model susceptibility to intentional linguistic adversarial attacks as well as natural linguistic variations is essential in this pursuit. We present extensive analysis of model robustness and susceptibility to linguistic variations in the setting of deceptive news detection, a difficult classification task that is an increasingly important problem to solve with the impact of misinformation spread online. We evaluate the effectiveness of incorporating adversarial defense strategies and measure model susceptibility to state-of-the-art adversarial attacks using two types of linguistic attacks — character and word perturbations. We consider two multiclass prediction tasks — a 3-way classification of tweets as trustworthy, propaganda, or disinformation; and a 4-way classification as clickbait, hoax, satire, or conspiracy — and compare the performance of three embeddings that have been state-of-the-art for several NLP tasks — GloVe, ELMo, and BERT — to highlight consistent trends in susceptibility, high confidence misclassifications, and high impact failures. We find that character or mixed ensemble models are the most effective defense mechanisms and that character perturbations are a more effective attack than word perturbations for deception classification.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Glenski, Maria F., Ayton, Ellyn M., Cosbey, Robin J., Arendt, Dustin L., Volkova, Svitlana. 2021-06-10. Evaluating Deception Detection Model Robustness To Linguistic Variation. https://doi.org/10.18653/v1%2F2021.socialnlp-1.6

Cite the original work for its findings. Save a collection to share your selection of sources.