BanFakeNews: A Dataset for Detecting Fake News in Bangla
arXiv:2004.08789
Abstract
Observing the damages that can be done by the rapid propagation of fake news in various sectors like politics and finance, automatic identification of fake news using linguistic analysis has drawn the attention of the research community. However, such methods are largely being developed for English where low resource languages remain out of the focus. But the risks spawned by fake and manipulative news are not confined by languages. In this work, we propose an annotated dataset of ~50K news that can be used for building automated fake news detection systems for a low resource language like Bangla. Additionally, we provide an analysis of the dataset and develop a benchmark system with state of the art NLP techniques to identify Bangla fake news. To create this system, we explore traditional linguistic features and neural network based methods. We expect this dataset will be a valuable resource for building technologies to prevent the spreading of fake news and contribute in research with low resource languages.
LREC 2020
References in corpus (3)
Cited by in corpus (8)
- Indonesia's Fake News Detection using Transformer Network
- Hostility Detection Dataset in Hindi
- Bangla Fake News Detection Based On Multichannel Combined CNN-LSTM
- BanglaBait: Semi-Supervised Adversarial Approach for Clickbait Detection on Bangla Clickbait Dataset
- Breaking the Fake News Barrier: Deep Learning Approaches in Bangla Language
- Improving Bangla Linguistics: Advanced LSTM, Bi-LSTM, and Seq2Seq Models for Translating Sylheti to Modern Bangla
- Ensuring the Inclusive Use of Natural Language Processing in the Global Response to COVID-19
- UPV at CheckThat! 2021: Mitigating Cultural Differences for Identifying Multilingual Check-worthy Claims