Large-scale analysis of Zipf's law in English texts
arXiv:1509.04486 · doi:10.1371/journal.pone.0147073
Abstract
Despite being a paradigm of quantitative linguistics, Zipf's law for words suffers from three main problems: its formulation is ambiguous, its validity has not been tested rigorously from a statistical point of view, and it has not been confronted to a representatively large number of texts. So, we can summarize the current support of Zipf's law in texts as anecdotic. We try to solve these issues by studying three different versions of Zipf's law and fitting them to all available English texts in the Project Gutenberg database (consisting of more than 30000 texts). To do so we use state-of-the art tools in fitting and goodness-of-fit tests, carefully tailored to the peculiarities of text statistics. Remarkably, one of the three versions of Zipf's law, consisting of a pure power-law form in the complementary cumulative distribution function of word frequencies, is able to fit more than 40% of the texts in the database (at the 0.05 significance level), for the whole domain of frequencies (from 1 to the maximum value) and with only one free parameter (the exponent).
References in corpus (12)
- Power-law distributions in empirical data
- Languages cool as they expand: Allometric scaling and the decreasing need for new words
- Parameter estimation for power-law distributions by maximum likelihood methods
- Understanding scaling through history-dependent processes with collapsing sample space
- Statistical laws in linguistics
- Zipf's law for word frequencies: word forms versus lemmas in long texts
- A maximum entropy framework for non-exponential distributions
- Log-log Convexity of Type-Token Growth in Zipf's Systems
- Text mixing shapes the anatomy of rank-frequency distributions: A modern Zipfian mechanics for natural language
- Statistical Patterns in Written Language
- A practical recipe to fit discrete power-law distributions
- Fitting and goodness-of-fit test of non-truncated and truncated power-law distributions
Cited by in corpus (30)
- Power-law size distributions in geoscience revisited
- True scale-free networks hidden by finite size effects
- Optimal coding and the origins of Zipfian laws
- Compression and the origins of Zipf's law for word frequencies
- Linguistic laws in biology
- The origins of Zipf's meaning-frequency law
- Word frequency and sentiment analysis of twitter messages during Coronavirus pandemic
- Testing statistical laws in complex systems
- Polysemy and brevity versus frequency in language
- The distinct flavors of Zipf's law in the rank-size and in the size-distribution representations, and its maximum-likelihood fitting
- The brevity law as a scaling law, and a possible origin of Zipf's law for word frequencies
- Generalized Entropies and the Similarity of Texts
- Maximal Diversity and Zipf's Law
- Truncated lognormal distributions and scaling in the size of naturally defined population clusters
- NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter Access
- On the emergence of Zipf's law in music
- Universality of power-law exponents by means of maximum likelihood estimation
- In Nomine Function: Naming Functions in Stripped Binaries with Neural Networks
- Lognormals, Power Laws and Double Power Laws in the Distribution of Frequencies of Harmonic Codewords from Classical Music
- Exploiting Data Skew for Improved Query Performance
- Heaps' Law and Vocabulary Richness in the History of Classical Music Harmony
- From Boltzmann to Zipf through Shannon and Jaynes
- A Power Law Approach to Estimating Fake Social Network Accounts
- Empirical Analysis of Zipf's Law, Power Law, and Lognormal Distributions in Medical Discharge Reports
- Zipf's law emerges asymptotically during phase transitions in communicative systems
- Universal and non-universal text statistics: Clustering coefficient for language identification
- Study of scaling laws in language families
- Quadratic Term Correction on Heaps' Law
- Distributional Ground Truth: Non-Redundant Crowdsourcing Data Quality Control in UI Labeling Tasks
- Modeling natural language emergence with integral transform theory and reinforcement learning