Rank-frequency relation for Chinese characters
arXiv:1309.1536 · doi:10.1140/epjb/e2014-40805-2
Abstract
We show that the Zipf's law for Chinese characters perfectly holds for sufficiently short texts (few thousand different characters). The scenario of its validity is similar to the Zipf's law for words in short English texts. For long Chinese texts (or for mixtures of short Chinese texts), rank-frequency relations for Chinese characters display a two-layer, hierarchic structure that combines a Zipfian power-law regime for frequent characters (first layer) with an exponential-like regime for less frequent characters (second layer). For these two layers we provide different (though related) theoretical descriptions that include the range of low-frequency characters (hapax legomena). The comparative analysis of rank-frequency relations for Chinese characters versus English words illustrates the extent to which the characters play for Chinese writers the same role as the words for those writing within alphabetical systems.
To appear in European Physical Journal B (EPJ B), 2014 (22 pages, 7 figures)
References in corpus (10)
- Power-law distributions in empirical data
- Probabilistic Latent Semantic Analysis
- Parameter estimation for power-law distributions by maximum likelihood methods
- Zipf's Law Leads to Heaps' Law: Analyzing Their Relation in Finite-Size Systems
- Zipf's Law and Avoidance of Excessive Synonymy
- Scaling Laws in Human Language
- A short account of a connection of Power Laws to the Information Entropy
- Explaining Zipf's Law via Mental Lexicon
- Size dependent word frequencies and translational invariance of books
- Maximum entropy approach to power-law distributions in coupled dynamic-stochastic systems
Cited by in corpus (7)
- Computational Socioeconomics
- On the emergence of Zipf's law in music
- Direct and indirect evidence of compression of word lengths. Zipf's law of abbreviation revisited
- Two halves of a meaningful text are statistically different
- The Dependence of Frequency Distributions on Multiple Meanings of Words, Codes and Signs
- Stochastic model for phonemes uncovers an author-dependency of their usage
- Non-random structures in universal compression and the Fermi paradox