On the origin of long-range correlations in texts
arXiv:1207.0658 · doi:10.1073/pnas.1117723109
Abstract
The complexity of human interactions with social and natural phenomena is mirrored in the way we describe our experiences through natural language. In order to retain and convey such a high dimensional information, the statistical properties of our linguistic output has to be highly correlated in time. An example are the robust observations, still largely not understood, of correlations on arbitrary long scales in literary texts. In this paper we explain how long-range correlations flow from highly structured linguistic levels down to the building blocks of a text (words, letters, etc..). By combining calculations and data analysis we show that correlations take form of a bursty sequence of events once we approach the semantically relevant topics of the text. The mechanisms we identify are fairly general and can be equally applied to other hierarchical settings.
Full paper (8 pages) and Supporting Information (19 pages)
References in corpus (2)
Cited by in corpus (10)
- Languages cool as they expand: Allometric scaling and the decreasing need for new words
- A joint text mining-rank size investigation of the rhetoric structures of the US Presidents' speeches
- Complexity-entropy analysis at different levels of organization in written language
- Tensor network language model
- Are queries and keys always relevant? A case study on Transformer wave functions
- Ordinal analysis of lexical patterns
- Entropy and type-token ratio in gigaword corpora
- Exploring language relations through syntactic distances and geographic proximity
- Correlation Dimension of Natural Language in a Statistical Manifold
- A Zipf-preserving, long-range correlated surrogate for written language and other symbolic sequences