FANG-COVID: A New Large-Scale Benchmark Dataset for Fake News Detection in German
Abstract
As the world continues to fight the COVID-19 pandemic, it is simultaneously fighting an 'infodemic' -a flood of disinformation and spread of conspiracy theories leading to health threats and the division of society.To combat this infodemic, there is an urgent need for benchmark datasets that can help researchers develop and evaluate models geared towards automatic detection of disinformation.While there are increasing efforts to create adequate, open-source benchmark datasets for English, comparable resources are virtually unavailable for German, leaving research for the German language lagging significantly behind.In this paper, we introduce the new benchmark dataset FANG-COVID consisting of 28,056 real and 13,186 fake German news articles related to the COVID-19 pandemic as well as data on their propagation on Twitter.Furthermore, we propose an explainable textual-and social context-based model for fake news detection, compare its performance to "blackbox" models and perform feature ablation to assess the relative importance of humaninterpretable features in distinguishing fake news from authentic news.
Authors 4
-
Justus Mattern Aachen
Affiliation as printed
RWTH Aachen University
-
Elma Kerz Aachen
Affiliation as printed
RWTH Aachen University
-
Affiliation as printed
University of Amsterdam
-
Markus Strohmaier Aachen
Affiliation as printed
RWTH Aachen University
Cited by 14 stored of 14
14 results
No patents citing this paper on Lens.org (checked 2026-10-06).
References 55
-
W2970597249details pending0citations
-
W2763572884details pending0citations
-
W1481519903details pending0citations
-
W1506509537details pending0citations
-
W2133130435details pending0citations
-
W2593408211details pending0citations
-
W2611921107details pending0citations
-
W2742330194details pending0citations