Statistical Patterns in Written Language
arXiv:1412.3336
Abstract
Quantitative linguistics has been allowed, in the last few decades, within the admittedly blurry boundaries of the field of complex systems. A growing host of applied mathematicians and statistical physicists devote their efforts to disclose regularities, correlations, patterns, and structural properties of language streams, using techniques borrowed from statistics and information theory. Overall, results can still be categorized as modest, but the prospects are promising: medium- and long-range features in the organization of human language -which are beyond the scope of traditional linguistics- have already emerged from this kind of analysis and continue to be reported, contributing a new perspective to our understanding of this most complex communication system. This short book is intended to review some of these recent contributions.
Some authors of work reviewed in this article have claimed rights on its graphical material. This material cannot be eliminated from the article without jeopardizing its coherence
References in corpus (2)
Cited by in corpus (9)
- Large-scale analysis of Zipf's law in English texts
- Zipf's law for word frequencies: word forms versus lemmas in long texts
- Statistical laws in linguistics
- Log-log Convexity of Type-Token Growth in Zipf's Systems
- Hierarchy of Scales in Language Dynamics
- Lognormals, Power Laws and Double Power Laws in the Distribution of Frequencies of Harmonic Codewords from Classical Music
- Heaps' Law and Vocabulary Richness in the History of Classical Music Harmony
- Variation of word frequencies in Russian literary texts
- Universal and non-universal text statistics: Clustering coefficient for language identification