Rank-frequency distribution of natural languages: a difference of probabilities approach
arXiv:1811.09451 · doi:10.1016/j.physa.2019.121795
Abstract
The time variation of the rank of words for six Indo-European languages is obtained using data from Google Books. For low ranks the distinct languages behave differently, maybe due to syntaxis rules, whereas for the law of large numbers predominates. The dynamics of is described stochastically through a master equation governing the time evolution of its probability density, which is approximated by a Fokker-Planck equation that is solved analytically. The difference between the data and the asymptotic solution is identified with the transient solution, and good agreement is obtained.
11 pages
References in corpus (5)
- The Matthew effect in empirical data
- Languages cool as they expand: Allometric scaling and the decreasing need for new words
- Evolution of the most common English words and phrases over the centuries
- Rank diversity of languages: Generic behavior in computational linguistics
- Robust clustering of languages across Wikipedia growth