Characterizing the Google Books corpus: Strong limits to inferences of socio-cultural and linguistic evolution
arXiv:1501.00960 · doi:10.1371/journal.pone.0137041
Abstract
It is tempting to treat frequency trends from the Google Books data sets as indicators of the "true" popularity of various words and phrases. Doing so allows us to draw quantitatively strong conclusions about the evolution of cultural perception of a given topic, such as time or gender. However, the Google Books corpus suffers from a number of limitations which make it an obscure mask of cultural popularity. A primary issue is that the corpus is in effect a library, containing one of each book. A single, prolific author is thereby able to noticeably insert new phrases into the Google Books lexicon, whether the author is widely read or not. With this understood, the Google Books corpus remains an important data set to be considered more lexicon-like than text-like. Here, we show that a distinct problematic feature arises from the inclusion of scientific texts, which have become an increasingly substantive portion of the corpus throughout the 1900s. The result is a surge of phrases typical to academic articles but less common in general, such as references to time in the form of citations. We highlight these dynamics by examining and comparing major contributions to the statistical divergence of English data sets between decades in the period 1800--2000. We find that only the English Fiction data set from the second version of the corpus is not heavily affected by professional texts, in clear contrast to the first version of the fiction data set and both unfiltered English data sets. Our findings emphasize the need to fully characterize the dynamics of the Google Books corpus before using these data sets to draw broad conclusions about cultural and linguistic evolution.
13 pages, 16 figures
References in corpus (1)
Cited by in corpus (40)
- The Geometry of Culture: Analyzing Meaning through Word Embeddings
- Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change
- Divergent discourse between protests and counter-protests: #BlackLivesMatter and #AllLivesMatter
- Mapping the Americanization of English in Space and Time
- Generalized Word Shift Graphs: A Method for Visualizing and Explaining Pairwise Comparisons Between Texts
- How we do things with words: Analyzing text as social and cultural data
- The Dynamics of Norm Change in the Cultural Evolution of Language
- A Hierarchy of Limitations in Machine Learning
- Storywrangler: A massive exploratorium for sociolinguistic, cultural, socioeconomic, and political timelines using Twitter
- Frequency patterns of semantic change: Corpus-based evidence of a near-critical dynamics in language change
- Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora
- Population size predicts lexical diversity, but so does the mean sea level - why it is important to correctly account for the structure of temporal data
- Benchmarking sentiment analysis methods for large-scale texts: A case for using continuum-scored words and word shift graphs
- Allotaxonometry and rank-turbulence divergence: A universal instrument for comparing complex systems
- English verb regularization in books and tweets
- Digital interfaces of historical newspapers: opportunities, restrictions and recommendations
- Postmortem memory of public figures in news and social media
- General Tensor Spectral Co-clustering for Higher-Order Data
- Hurricanes and hashtags: Characterizing online collective attention for natural disasters
- Challenges in detecting evolutionary forces in language change using diachronic corpora
- Challenges for Computational Lexical Semantic Change
- The Natural Selection of Words: Finding the Features of Fitness
- Intellectual interchanges in the history of the massive online open-editing encyclopedia, Wikipedia
- Reply to Garcia et al.: Common mistakes in measuring frequency dependent word characteristics
- Computational Paremiology: Charting the temporal, ecological dynamics of proverb use in books, news articles, and tweets
- Characterizing English Variation across Social Media Communities with BERT
- Communicative need modulates competition in language change
- A decomposition of book structure through ousiometric fluctuations in cumulative word-time
- Computational Sociolinguistics: A Survey
- Towards a science of human stories: using sentiment analysis and emotional arcs to understand the building blocks of complex social systems
- History Playground: A Tool for Discovering Temporal Trends in Massive Textual Corpora
- Enhancing Task-Oriented Dialogues with Chitchat: a Comparative Study Based on Lexical Diversity and Divergence
- Ousiometrics: The essence of meaning aligns with a power-danger-structure framework instead of valence-arousal-dominance
- Subdiffusive semantic evolution in Indo-European languages
- Towards Evaluation of Cultural-scale Claims in Light of Topic Model Sampling Effects
- Fake news as we feel it: perception and conceptualization of the term "fake news" in the media
- A Statistical Model of Word Rank Evolution
- A model of phase-coupled delay equations for the dynamics of word usage
- Cognitive forces shape the dynamics of word usage across multiple languages
- Oscillatory dynamics between language usage and economic activity