Dataset information
Available languages
English
Dataset description
This is the dataset presented in the paper 'Mute Cods: A multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection', accepted at LREC conference 2026, by Laken, Marino, Piot, Bassi, Fomsgaard, Maggini, Vieira, García and Tonelli. The dataset consists of 5750 messages across English, Dutch, Italian, Spanish and Portuguese from 87 channels documented as disseminating conspiracist and extremist content. Domain experts annotated messages for conspiracist tone, population replacement conspiracy theories, vaccine conspiracies, and hate speech. This dataset is pseudonomyzed. We release both the raw annotations and the aggregated labels with train/test split used to train the models as reported in the paper. As this is social media user data, we request users not to share the dataset with third parties; rather, send them the link to our repository, and we will grant them access as well. If you use this dataset, please cite our paper: Katarina Laken, Erik Bran Marino, Paloma Piot, Davide Bassi, Søren Fomsgaard, Michele Maggini, Renata Vieira, Marcos García, and Sara Tonelli. 2026. Mute Cods: A Multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection. In Proceedings of the Fifteenth International Conference on Language Resources and Evaluation (LREC 2026). Palma de Mallorca, Spain.
European data infrastructure with broad catalog discovery, free evaluation access and production-grade API options.
190K+
indexed dataset pages
32
countries and EU institutions
Free API quota
for evaluation and prototypes
SLA
history and push on production APIs