Storywrangler: A massive exploratorium for sociolinguistic, cultural, socioeconomic, and political timelines using Twitter
arXiv:2007.12988 · doi:10.1126/sciadv.abe6534
Abstract
In real-time, social media data strongly imprints world events, popular culture, and day-to-day conversations by millions of ordinary people at a scale that is scarcely conventionalized and recorded. Vitally, and absent from many standard corpora such as books and news archives, sharing and commenting mechanisms are native to social media platforms, enabling us to quantify social amplification (i.e., popularity) of trending storylines and contemporary cultural phenomena. Here, we describe Storywrangler, a natural language processing instrument designed to carry out an ongoing, day-scale curation of over 100 billion tweets containing roughly 1 trillion 1-grams from 2008 to 2021. For each day, we break tweets into unigrams, bigrams, and trigrams spanning over 100 languages. We track n-gram usage frequencies, and generate Zipf distributions, for words, hashtags, handles, numerals, symbols, and emojis. We make the data set available through an interactive time series viewer, and as downloadable time series and daily distributions. Although Storywrangler leverages Twitter data, our method of extracting and tracking dynamic changes of n-grams can be extended to any similar social media platform. We showcase a few examples of the many possible avenues of study we aim to enable including how social amplification can be visualized through 'contagiograms'. We also present some example case studies that bridge n-gram time series with disparate data sources to explore sociotechnical dynamics of famous individuals, box office success, and social unrest.
Main text: 15 pages, 6 figures; Supplementary text: 23 pages, 11 figures, 15 tables. Website: https://storywrangling.org/
References in corpus (3)
- How the world's collective attention is being paid to a pandemic: COVID-19 related n-gram time series for 24 languages on Twitter
- The growing amplification of social media: Measuring temporal and social contagion dynamics for over 150 languages on Twitter for 2009-2020
- Computational timeline reconstruction of the stories surrounding Trump: Story turbulence, narrative control, and collective chronopathy
Cited by in corpus (14)
- How the world's collective attention is being paid to a pandemic: COVID-19 related n-gram time series for 24 languages on Twitter
- The growing amplification of social media: Measuring temporal and social contagion dynamics for over 150 languages on Twitter for 2009-2020
- Sentiment and structure in word co-occurrence networks on Twitter
- American cultural regions mapped through the lexical analysis of social media
- Hurricanes and hashtags: Characterizing online collective attention for natural disasters
- Computational timeline reconstruction of the stories surrounding Trump: Story turbulence, narrative control, and collective chronopathy
- Augmenting semantic lexicons using word embeddings and transfer learning
- Entropy and type-token ratio in gigaword corpora
- When Dialects Collide: How Socioeconomic Mixing Affects Language Use
- Computational Paremiology: Charting the temporal, ecological dynamics of proverb use in books, news articles, and tweets
- Quantifying language changes surrounding mental health on Twitter
- Blending search queries with social media data to improve forecasts of economic indicators
- Language statistics at different spatial, temporal, and grammatical scales
- Probability-turbulence divergence: A tunable allotaxonometric instrument for comparing heavy-tailed categorical distributions