Statistical laws in linguistics
arXiv:1502.03296 · doi:10.1007/978-3-319-24403-7_2
Abstract
Zipf's law is just one out of many universal laws proposed to describe statistical regularities in language. Here we review and critically discuss how these laws can be statistically interpreted, fitted, and tested (falsified). The modern availability of large databases of written text allows for tests with an unprecedent statistical accuracy and also a characterization of the fluctuations around the typical behavior. We find that fluctuations are usually much larger than expected based on simplifying statistical assumptions (e.g., independence and lack of correlations between observations).These simplifications appear also in usual statistical tests so that the large fluctuations can be erroneously interpreted as a falsification of the law. Instead, here we argue that linguistic laws are only meaningful (falsifiable) if accompanied by a model for which the fluctuations can be computed (e.g., a generative model of the text). The large fluctuations we report show that the constraints imposed by linguistic laws on the creativity process of text generation are not as tight as one could expect.
Proceedings of the Flow Machines Workshop: Creativity and Universality in Language, Paris, June 18 to 20, 2014
References in corpus (15)
- Power-law distributions in empirical data
- Fluctuation scaling in complex systems: Taylor's law and beyond
- Beyond word frequency: Bursts, lulls, and scaling in the temporal distributions of words
- Languages cool as they expand: Allometric scaling and the decreasing need for new words
- Parameter estimation for power-law distributions by maximum likelihood methods
- Zipf's Law Leads to Heaps' Law: Analyzing Their Relation in Finite-Size Systems
- On the origin of long-range correlations in texts
- Scaling: Lost in the smog
- Scaling laws and fluctuations in the statistics of word frequencies
- A maximum entropy framework for non-exponential distributions
- Optimization models of natural communication
- Text mixing shapes the anatomy of rank-frequency distributions: A modern Zipfian mechanics for natural language
- Statistical Patterns in Written Language
- Size dependent word frequencies and translational invariance of books
- Fitting and goodness-of-fit test of non-truncated and truncated power-law distributions
Cited by in corpus (23)
- A network approach to topic models
- Large-scale analysis of Zipf's law in English texts
- Linguistic laws in biology
- Testing statistical laws in complex systems
- Complex systems approach to natural language
- The distinct flavors of Zipf's law in the rank-size and in the size-distribution representations, and its maximum-likelihood fitting
- Zipf and Heaps laws from dependency structures in component systems
- The brevity law as a scaling law, and a possible origin of Zipf's law for word frequencies
- Generalized Entropies and the Similarity of Texts
- Multilayer Networks for Text Analysis with Multiple Data Types
- On the emergence of Zipf's law in music
- Heaps' law, statistics of shared components and temporal patterns from a sample-space-reducing process
- Lognormals, Power Laws and Double Power Laws in the Distribution of Frequencies of Harmonic Codewords from Classical Music
- A note on retrodiction and machine evolution
- Compression and the origins of Zipf's law of abbreviation
- Heaps' Law and Vocabulary Richness in the History of Classical Music Harmony
- Assessing Language Models with Scaling Properties
- Variation of word frequencies in Russian literary texts
- Towards Controllable and Personalized Review Generation
- Universal and non-universal text statistics: Clustering coefficient for language identification
- Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
- Computational lexical analysis of Flamenco genres
- Evaluating Computational Language Models with Scaling Properties of Natural Language