Zipf's law unzipped
arXiv:1104.1789 · doi:10.1088/1367-2630/13/4/043004
Abstract
Why does Zipf's law give a good description of data from seemingly completely unrelated phenomena? Here it is argued that the reason is that they can all be described as outcomes of a ubiquitous random group division: the elements can be citizens of a country and the groups family names, or the elements can be all the words making up a novel and the groups the unique words, or the elements could be inhabitants and the groups the cities in a country, and so on. A Random Group Formation (RGF) is presented from which a Bayesian estimate is obtained based on minimal information: it provides the best prediction for the number of groups with elements, given the total number of elements, groups, and the number of elements in the largest group. For each specification of these three values, the RGF predicts a unique group distribution , where the power-law index is a unique function of the same three values. The universality of the result is made possible by the fact that no system specific assumptions are made about the mechanism responsible for the group division. The direct relation between and the total number of elements, groups, and the number of elements in the largest group, is calculated. The predictive power of the RGF model is demonstrated by direct comparison with data from a variety of systems. It is shown that usually takes values in the interval and that the value for a given phenomena depends in a systematic way on the total size of the data set. The results are put in the context of earlier discussions on Zipf's and Gibrat's laws, and the connection between growth models and RGF is elucidated.
22 pages, 32 figures
References in corpus (3)
Cited by in corpus (47)
- Colloquium: Criticality and dynamical scaling in living systems
- The Matthew effect in empirical data
- Languages cool as they expand: Allometric scaling and the decreasing need for new words
- Stochastic model for the vocabulary growth in natural languages
- Zipf's law in 50 languages: its structural pattern, linguistic interpretation, and cognitive motivation
- Emotional persistence in online chatting communities
- Social Contagion: An Empirical Study of Information Spread on Digg and Twitter Follower Graphs
- Optimal coding and the origins of Zipfian laws
- Distance weighted city growth
- Similarity of symbol frequency distributions with heavy tails
- Rank diversity of languages: Generic behavior in computational linguistics
- The workings of the Maximum Entropy Principle in collective human behavior
- Zipf's Law from Scale-free Geometry
- Statistics of shared components in complex component systems
- Space-time correlations in urban sprawl
- Zipf's and Taylor's Laws
- Rank-frequency relation for Chinese characters
- MaxEnt and dynamical information
- Explaining Zipf's Law via Mental Lexicon
- Principle of least effort vs maximum efficiency: deriving Zipf-Pareto laws
- On the emergence of Zipf's law in music
- Universal scaling in sports ranking
- Rank dynamics of word usage at multiple scales
- Heaps' law, statistics of shared components and temporal patterns from a sample-space-reducing process
- A Paradoxical Property of the Monkey Book
- The Ten Thousand Kims
- Memory endowed US cities and their demographic interactions
- Benford's Law and First Letter of Word
- Maximum Entropy, Word-Frequency, Chinese Characters, and Multiple Meanings
- Multifractal analysis of sentence lengths in English literary texts
- Randomness versus specifics for word-frequency distributions
- Two halves of a meaningful text are statistically different
- The Dependence of Frequency Distributions on Multiple Meanings of Words, Codes and Signs
- Family of Probability Distributions Derived from Maximal Entropy Principle with Scale Invariant Restrictions
- The likely determines the unlikely
- Growth or Reproduction: Emergence of an Evolutionary Optimal Strategy
- Universal statistics of the knockout tournament
- Language statistics at different spatial, temporal, and grammatical scales
- ScienceWISE: Topic Modeling over Scientific Literature Networks
- Co-occurrence of the Benford-like and Zipf Laws Arising from the Texts Representing Human and Artificial Languages
- Complex distributions emerging in filtering and compression
- Filtering Statistics on Networks
- Surname statistics - Crossing the boundary between disciplines
- Why Money Trickles Up - Wealth & Income Distributions
- Computer activity learning from system call time series
- Influence Process Structural Learning and the Emergence of Collective Intelligence
- Dynamics of Users Activity on Web-Blogs