Entropy and Long range correlations in literary English
arXiv:cond-mat/0204108 · doi:10.1209/0295-5075/26/4/001
Abstract
We investigated long range correlations in two literary texts, Moby Dick by H. Melville and Grimm's tales. The analysis is based on the calculation of entropy like quantities as the mutual information for pairs of letters and the entropy, the mean uncertainty, per letter. We further estimate the number of different subwords of a given length n. Filtering out the contributions due to the effects of the finite length of the texts, we find correlations ranging to a few hundred letters. Scaling laws for the mutual information (decay with a power law), for the entropy per letter (decay with the inverse square root of n) and for the word numbers (stretched exponential growth with n and with a power law of the text length) were found.
6 pages, 3 figures
Cited by in corpus (53)
- Scaling Laws for Neural Language Models
- Entropy estimation of symbol sequences
- Zipf's Law Leads to Heaps' Law: Analyzing Their Relation in Finite-Size Systems
- Machine Learning for Quantum Matter
- On the origin of long-range correlations in texts
- Time's Barbed Arrow: Irreversibility, Crypticity, and Stored Information
- Prediction, Retrodiction, and The Amount of Information Stored in the Present
- Statistical laws in linguistics
- Statistical keyword detection in literary corpora
- Complexity Through Nonextensivity
- Information content versus word length in random typing
- Scaling Laws in Human Language
- Towards the quantification of the semantic information encoded in written language
- Testing reanalysis datasets in Antarctica: Trends, persistence properties and trend significance
- The Mathematical Relationship between Zipf's Law and the Hierarchical Scaling Law
- Criticality in Formal Languages and Statistical Physics
- On the Vocabulary of Grammar-Based Codes and the Logical Consistency of Texts
- The span of correlations in dolphin whistle sequences
- Random Language Model
- Guessing probability distributions from small samples
- Excess entropy in natural language: present state and perspectives
- Statistical Patterns in Written Language
- Entropy and hierarchical clustering: characterising the morphology of the urban fabric in different spatial cultures
- Is Natural Language a Perigraphic Process? The Theorem about Facts and Words Revisited
- Constant conditional entropy and related hypotheses
- On Hidden Markov Processes with Infinite Excess Entropy
- Do Neural Nets Learn Statistical Laws behind Natural Language?
- Fractal Power Law in Literary English
- On Hilberg's Law and Its Links with Guiraud's Law
- Can coarse-graining introduce long-range correlations in a symbolic sequence?
- Tensor network language model
- Mutual Information Scaling and Expressive Power of Sequence Models
- Understanding Recurrent Neural Architectures by Analyzing and Synthesizing Long Distance Dependencies in Benchmark Sequential Datasets
- The quoter model: a paradigmatic model of the social flow of written information
- Multifractal analysis of sentence lengths in English literary texts
- Pull out all the stops: Textual analysis via punctuation sequences
- Ordinal analysis of lexical patterns
- Information-theoretical analysis of the statistical dependencies among three variables: Applications to written language
- Entropy and type-token ratio in gigaword corpora
- Assessing Language Models with Scaling Properties
- The Past and the Future in the Present
- Trimming the Independent Fat: Sufficient Statistics, Mutual Information, and Predictability from Effective Channel States
- Parallels of human language in the behavior of bottlenose dolphins
- From form to information: Analysing built environments in different spatial cultures
- Unsupervised detection of semantic correlations in big data
- Taylor's law for Human Linguistic Sequences
- Segmentation and Context of Literary and Musical Sequences
- A Preadapted Universal Switch Distribution for Testing Hilberg's Conjecture
- Evaluating Computational Language Models with Scaling Properties of Natural Language
- Long-range correlations and trends in Colombian seismic time series
- Correction algorithm for finite sample statistics
- Bounds for Algorithmic Mutual Information and a Unifilar Order Estimator
- A Zipf-preserving, long-range correlated surrogate for written language and other symbolic sequences