Stochastic model for the vocabulary growth in natural languages
arXiv:1212.1362 · doi:10.1103/PhysRevX.3.021006
Abstract
We propose a stochastic model for the number of different words in a given database which incorporates the dependence on the database size and historical changes. The main feature of our model is the existence of two different classes of words: (i) a finite number of core-words which have higher frequency and do not affect the probability of a new word to be used; and (ii) the remaining virtually infinite number of noncore-words which have lower frequency and once used reduce the probability of a new word to be used in the future. Our model relies on a careful analysis of the google-ngram database of books published in the last centuries and its main consequence is the generalization of Zipf's and Heaps' law to two scaling regimes. We confirm that these generalizations yield the best simple description of the data among generic descriptive models and that the two free parameters depend only on the language but not on the database. From the point of view of our model the main change on historical time scales is the composition of the specific words included in the finite list of core-words, which we observe to decay exponentially in time with a rate of approximately 30 words per year for English.
corrected typos and errors in reference list; 10 pages text, 15 pages supplemental material; to appear in Physical Review X
References in corpus (5)
- Statistical physics of social dynamics
- Languages cool as they expand: Allometric scaling and the decreasing need for new words
- Collective dynamics of social annotation
- Culturomics meets random fractal theory: Insights into long-range correlations of social and natural phenomena over the past two centuries
- Wikipedia information flow analysis reveals the scale-free architecture of the Semantic Space
Cited by in corpus (55)
- Characterizing the Google Books corpus: Strong limits to inferences of socio-cultural and linguistic evolution
- A network approach to topic models
- Network dynamics of innovation processes
- Mapping the Americanization of English in Space and Time
- Large-scale analysis of Zipf's law in English texts
- Computational Socioeconomics
- Zipf's law in 50 languages: its structural pattern, linguistic interpretation, and cognitive motivation
- Internal and external dynamics in language: Evidence from verb regularity in a historical corpus of English
- Natural Language Statistical Features of LSTM-generated Texts
- Zipf's law for word frequencies: word forms versus lemmas in long texts
- Statistical laws in linguistics
- Scaling laws and fluctuations in the statistics of word frequencies
- Evolution and structure of technological systems - An innovation output network
- Testing statistical laws in complex systems
- A scaling law beyond Zipf's law and its relation to Heaps' law
- Similarity of symbol frequency distributions with heavy tails
- Optimization models of natural communication
- Rank diversity of languages: Generic behavior in computational linguistics
- In narrative texts punctuation marks obey the same statistics as words
- The distinct flavors of Zipf's law in the rank-size and in the size-distribution representations, and its maximum-likelihood fitting
- Log-log Convexity of Type-Token Growth in Zipf's Systems
- Text mixing shapes the anatomy of rank-frequency distributions: A modern Zipfian mechanics for natural language
- Zipf and Heaps laws from dependency structures in component systems
- Complexity measurement of natural and artificial languages
- Generalized Entropies and the Similarity of Texts
- Spatial evolution of human dialects
- Statistics of shared components in complex component systems
- Truncated lognormal distributions and scaling in the size of naturally defined population clusters
- Innovation and Nested Preferential Growth in Chess Playing Behavior
- Rank dynamics of word usage at multiple scales
- Heaps' law, statistics of shared components and temporal patterns from a sample-space-reducing process
- Simon's fundamental rich-get-richer model entails a dominant first-mover advantage
- Empirical observations of ultraslow diffusion driven by the fractional dynamics in languages: Dynamical statistical properties of word counts of already popular words
- Scaling laws and dynamics of hashtags on Twitter
- The meaning-frequency law in Zipfian optimization models of communication
- Lognormals, Power Laws and Double Power Laws in the Distribution of Frequencies of Harmonic Codewords from Classical Music
- Compression and the origins of Zipf's law of abbreviation
- Stochastic dynamics and the predictability of big hits in online videos
- Beyond the Chinese Restaurant and Pitman-Yor processes: Statistical Models with Double Power-law Behavior
- The infochemical core
- Quantifying the Dissimilarity of Texts
- Entropy and type-token ratio in gigaword corpora
- Information-theoretical analysis of the statistical dependencies among three variables: Applications to written language
- Damping effect in innovation processes: case studies from Twitter
- Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models
- Variation of word frequencies in Russian literary texts
- Language statistics at different spatial, temporal, and grammatical scales
- Word frequency-rank relationship in tagged texts
- A minor extension of the logistic equation for growth of word counts on online media: Parametric description of diversity of growth phenomena in society
- Co-occurrence of the Benford-like and Zipf Laws Arising from the Texts Representing Human and Artificial Languages
- Markovian language model of the DNA and its information content
- Evaluating Computational Language Models with Scaling Properties of Natural Language
- A general solution to the preferential selection model
- Zipf's laws of meaning in Catalan
- Verifying Heaps' law using Google Books Ngram data