Complexity measurement of natural and artificial languages
arXiv:1311.5427 · doi:10.1002/cplx.21529
Abstract
We compared entropy for texts written in natural languages (English, Spanish) and artificial languages (computer software) based on a simple expression for the entropy as a function of message length and specific word diversity. Code text written in artificial languages showed higher entropy than text of similar length expressed in natural languages. Spanish texts exhibit more symbolic diversity than English ones. Results showed that algorithms based on complexity measures differentiate artificial from natural languages, and that text analysis based on complexity measures allows the unveiling of important aspects of their nature. We propose specific expressions to examine entropy related aspects of tests and estimate the values of entropy, emergence, self-organization and complexity based on specific diversity and message length.
29 pages, 11 figures, 3 tables, 2 appendixes
References in corpus (2)
Cited by in corpus (10)
- Statistical laws in linguistics
- Measuring the Complexity of Self-organizing Traffic Lights
- Rank diversity of languages: Generic behavior in computational linguistics
- Quantifying literature quality using complexity criteria
- Emergence in artificial life
- Rank dynamics of word usage at multiple scales
- Measuring the Complexity of Continuous Distributions
- Calculating entropy at different scales among diverse communication systems
- A Fundamental Scale of Descriptions for Analyzing Information Content of Communication Systems
- Relating complexities for the reflexive study of complex systems