Text mixing shapes the anatomy of rank-frequency distributions: A modern Zipfian mechanics for natural language
arXiv:1409.3870 · doi:10.1103/PhysRevE.91.052811
Abstract
Natural languages are full of rules and exceptions. One of the most famous quantitative rules is Zipf's law which states that the frequency of occurrence of a word is approximately inversely proportional to its rank. Though this `law' of ranks has been found to hold across disparate texts and forms of data, analyses of increasingly large corpora over the last 15 years have revealed the existence of two scaling regimes. These regimes have thus far been explained by a hypothesis suggesting a separability of languages into core and non-core lexica. Here, we present and defend an alternative hypothesis, that the two scaling regimes result from the act of aggregating texts. We observe that text mixing leads to an effective decay of word introduction, which we show provides accurate predictions of the location and severity of breaks in scaling. Upon examining large corpora from 10 languages in the Project Gutenberg eBooks collection (eBooks), we find emphatic empirical support for the universality of our claim.
9 pages, 6 figures, and 1 table
References in corpus (1)
Cited by in corpus (9)
- Generalized Entropies and the Similarity of Texts
- Sentiment and structure in word co-occurrence networks on Twitter
- Information flow estimation: a study of news on Twitter
- Lognormals, Power Laws and Double Power Laws in the Distribution of Frequencies of Harmonic Codewords from Classical Music
- Reply to Garcia et al.: Common mistakes in measuring frequency dependent word characteristics
- Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models
- Probability-turbulence divergence: A tunable allotaxonometric instrument for comparing heavy-tailed categorical distributions
- Complete asymptotic type-token relationship for growing complex systems with inverse power-law count rankings
- Zipf's laws of meaning in Catalan