Explaining Zipf's Law via Mental Lexicon
arXiv:1302.4383 · doi:10.1103/PhysRevE.88.062804
Abstract
The Zipf's law is the major regularity of statistical linguistics that served as a prototype for rank-frequency relations and scaling laws in natural sciences. Here we show that the Zipf's law -- together with its applicability for a single text and its generalizations to high and low frequencies including hapax legomena -- can be derived from assuming that the words are drawn into the text with random probabilities. Their apriori density relates, via the Bayesian statistics, to general features of the mental lexicon of the author who produced the text.
References in corpus (6)
- Probabilistic Latent Semantic Analysis
- Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence (1999)
- Zipf's Law and Avoidance of Excessive Synonymy
- A short account of a connection of Power Laws to the Information Entropy
- Maximum entropy approach to power-law distributions in coupled dynamic-stochastic systems
- Analysis of an information-theoretic model for communication
Cited by in corpus (8)
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Rank-frequency relation for Chinese characters
- A comparative analysis of knowledge acquisition performance in complex networks
- Two halves of a meaningful text are statistically different
- Stochastic model for phonemes uncovers an author-dependency of their usage
- Markov Chain Monte Carlo for generating ranked textual data
- Non-random structures in universal compression and the Fermi paradox
- Unsupervised extraction of local and global keywords from a single text