How the world's collective attention is being paid to a pandemic: COVID-19 related n-gram time series for 24 languages on Twitter
arXiv:2003.12614 · doi:10.1371/journal.pone.0244476
Abstract
In confronting the global spread of the coronavirus disease COVID-19 pandemic we must have coordinated medical, operational, and political responses. In all efforts, data is crucial. Fundamentally, and in the possible absence of a vaccine for 12 to 18 months, we need universal, well-documented testing for both the presence of the disease as well as confirmed recovery through serological tests for antibodies, and we need to track major socioeconomic indices. But we also need auxiliary data of all kinds, including data related to how populations are talking about the unfolding pandemic through news and stories. To in part help on the social media side, we curate a set of 2000 day-scale time series of 1- and 2-grams across 24 languages on Twitter that are most 'important' for April 2020 with respect to April 2019. We determine importance through our allotaxonometric instrument, rank-turbulence divergence. We make some basic observations about some of the time series, including a comparison to numbers of confirmed deaths due to COVID-19 over time. We broadly observe across all languages a peak for the language-specific word for 'virus' in January 2020 followed by a decline through February and then a surge through March and April. The world's collective attention dropped away while the virus spread out from China. We host the time series on Gitlab, updating them on a daily basis while relevant. Our main intent is for other researchers to use these time series to enhance whatever analyses that may be of use during the pandemic as well as for retrospective investigations.
13 pages, 6 figures, 3 tables, website: http://compstorylab.org/covid19ngrams/
References in corpus (6)
- The COVID-19 Social Media Infodemic
- Tracking Social Media Discourse About the COVID-19 Pandemic: Development of a Public Coronavirus Twitter Data Set
- Tracking COVID-19 using online search
- The growing amplification of social media: Measuring temporal and social contagion dynamics for over 150 languages on Twitter for 2009-2020
- Storywrangler: A massive exploratorium for sociolinguistic, cultural, socioeconomic, and political timelines using Twitter
- Allotaxonometry and rank-turbulence divergence: A universal instrument for comparing complex systems
Cited by in corpus (12)
- Storywrangler: A massive exploratorium for sociolinguistic, cultural, socioeconomic, and political timelines using Twitter
- Causal Modeling of Twitter Activity During COVID-19
- Attention dynamics on the Chinese social media Sina Weibo during the COVID-19 pandemic
- ArCOV-19: The First Arabic COVID-19 Twitter Dataset with Propagation Networks
- Analyzing COVID-19 on Online Social Media: Trends, Sentiments and Emotions
- Large-scale, Language-agnostic Discourse Classification of Tweets During COVID-19
- Hurricanes and hashtags: Characterizing online collective attention for natural disasters
- Augmenting semantic lexicons using word embeddings and transfer learning
- Characterizing Twitter users behaviour during the Spanish Covid-19 first wave
- Long-term word frequency dynamics derived from Twitter are corrupted: A bespoke approach to detecting and removing pathologies in ensembles of time series
- Ousiometrics: The essence of meaning aligns with a power-danger-structure framework instead of valence-arousal-dominance
- An Exploration of Geo-temporal Characteristics of Users' Reactions on Social Media During the Pandemic