Eating Garlic Prevents COVID-19 Infection: Detecting Misinformation on the Arabic Content of Twitter
arXiv:2101.05626
Abstract
The rapid growth of social media content during the current pandemic provides useful tools for disseminating information which has also become a root for misinformation. Therefore, there is an urgent need for fact-checking and effective techniques for detecting misinformation in social media. In this work, we study the misinformation in the Arabic content of Twitter. We construct a large Arabic dataset related to COVID-19 misinformation and gold-annotate the tweets into two categories: misinformation or not. Then, we apply eight different traditional and deep machine learning models, with different features including word embeddings and word frequency. The word embedding models (\textsc{FastText} and word2vec) exploit more than two million Arabic tweets related to COVID-19. Experiments show that optimizing the area under the curve (AUC) improves the models' performance and the Extreme Gradient Boosting (XGBoost) presents the highest accuracy in detecting COVID-19 misinformation online.
18 pages, 4 figures
References in corpus (5)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Analysis of misinformation during the COVID-19 outbreak in China: cultural, social and political entanglements
- Cultural Convergence: Insights into the behavior of misinformation networks on Twitter
- ArCorona: Analyzing Arabic Tweets in the Early Days of Coronavirus (COVID-19) Pandemic
- Challenges in Combating COVID-19 Infodemic -- Data, Tools, and Ethics