Excess entropy in natural language: present state and perspectives
arXiv:1105.1306 · doi:10.1063/1.3630929
Abstract
We review recent progress in understanding the meaning of mutual information in natural language. Let us define words in a text as strings that occur sufficiently often. In a few previous papers, we have shown that a power-law distribution for so defined words (a.k.a. Herdan's law) is obeyed if there is a similar power-law growth of (algorithmic) mutual information between adjacent portions of texts of increasing length. Moreover, the power-law growth of information holds if texts describe a complicated infinite (algorithmically) random object in a highly repetitive way, according to an analogous power-law distribution. The described object may be immutable (like a mathematical or physical constant) or may evolve slowly in time (like cultural heritage). Here we reflect on the respective mathematical results in a less technical way. We also discuss feasibility of deciding to what extent these results apply to the actual human communication.
12 pages; no figures
References in corpus (3)
Cited by in corpus (8)
- Is Natural Language a Perigraphic Process? The Theorem about Facts and Words Revisited
- Divergent Predictive States: The Statistical Complexity Dimension of Stationary, Ergodic Hidden Markov Processes
- Constant conditional entropy and related hypotheses
- On Hidden Markov Processes with Infinite Excess Entropy
- Complexity-entropy analysis at different levels of organization in written language
- Natural Complexity, Computational Complexity and Depth
- Two halves of a meaningful text are statistically different
- Circumventing the Curse of Dimensionality in Prediction: Causal Rate-Distortion for Infinite-Order Markov Processes