A large-scale COVID-19 Twitter chatter dataset for open scientific research -- an international collaboration
arXiv:2004.03688 · doi:10.3390/epidemiologia2030024
Abstract
As the COVID-19 pandemic continues its march around the world, an unprecedented amount of open data is being generated for genetics and epidemiological research. The unparalleled rate at which many research groups around the world are releasing data and publications on the ongoing pandemic is allowing other scientists to learn from local experiences and data generated in the front lines of the COVID-19 pandemic. However, there is a need to integrate additional data sources that map and measure the role of social dynamics of such a unique world-wide event into biomedical, biological, and epidemiological analyses. For this purpose, we present a large-scale curated dataset of over 152 million tweets, growing daily, related to COVID-19 chatter generated from January 1st to April 4th at the time of writing. This open dataset will allow researchers to conduct a number of research projects relating to the emotional and mental responses to social distancing measures, the identification of sources of misinformation, and the stratified measurement of sentiment towards the pandemic in near real time.
8 pages, 1 figure 2 table. Update: new version of paper with up-to-date statistics and new co-authors
References in corpus (4)
Cited by in corpus (29)
- A large-scale COVID-19 Twitter chatter dataset for open scientific research -- an international collaboration
- Automated Detection and Forecasting of COVID-19 using Deep Learning Techniques: A Review
- TweetsCOV19 -- A Knowledge Base of Semantically Annotated Tweets about the COVID-19 Pandemic
- Public risk perception and emotion on Twitter during the Covid-19 pandemic
- Classification Aware Neural Topic Model and its Application on a New COVID-19 Disinformation Corpus
- TBCOV: Two Billion Multilingual COVID-19 Tweets with Sentiment, Entity, Geo, and Gender Labels
- ArCOV-19: The First Arabic COVID-19 Twitter Dataset with Propagation Networks
- Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset
- Is Working From Home The New Norm? An Observational Study Based on a Large Geo-tagged COVID-19 Twitter Dataset
- A Survey of COVID-19 Misinformation: Datasets, Detection Techniques and Open Issues
- AraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News and Hate Speech Detection Dataset
- Insight from NLP Analysis: COVID-19 Vaccines Sentiments on Social Media
- BillionCOV: An Enriched Billion-scale Collection of COVID-19 tweets for Efficient Hydration
- ArCorona: Analyzing Arabic Tweets in the Early Days of Coronavirus (COVID-19) Pandemic
- On Analyzing Antisocial Behaviors Amid COVID-19 Pandemic
- Challenges in Combating COVID-19 Infodemic -- Data, Tools, and Ethics
- Face Off: Polarized Public Opinions on Personal Face Mask Usage during the COVID-19 Pandemic
- Extracting a Knowledge Base of COVID-19 Events from Social Media
- Don't be a Victim During a Pandemic! Analysing Security and Privacy Threats in Twitter During COVID-19
- How Have We Reacted To The COVID-19 Pandemic? Analyzing Changing Indian Emotions Through The Lens of Twitter
- LEXpander: applying colexification networks to automated lexicon expansion
- Data and models for stance and premise detection in COVID-19 tweets: insights from the Social Media Mining for Health (SMM4H) 2022 shared task
- Global Tweet Mentions of COVID-19
- A Large-Scale Dataset of Search Interests Related to Disease X Originating from Different Geographic Regions
- TEST_POSITIVE at W-NUT 2020 Shared Task-3: Joint Event Multi-task Learning for Slot Filling in Noisy Text
- Unsupervised Text Mining of COVID-19 Records
- Findings of the NLP4IF-2021 Shared Tasks on Fighting the COVID-19 Infodemic and Censorship Detection
- Diagnosis of COVID-19 and Non-COVID-19 Patients by Classifying Only a Single Cough Sound
- EPIC30M: An Epidemics Corpus Of Over 30 Million Relevant Tweets