Probing the statistical properties of unknown texts: application to the Voynich Manuscript
arXiv:1303.0347 · doi:10.1371/journal.pone.0067310
Abstract
While the use of statistical physics methods to analyze large corpora has been useful to unveil many patterns in texts, no comprehensive investigation has been performed investigating the properties of statistical measurements across different languages and texts. In this study we propose a framework that aims at determining if a text is compatible with a natural language and which languages are closest to it, without any knowledge of the meaning of the words. The approach is based on three types of statistical measurements, i.e. obtained from first-order statistics of word properties in a text, from the topology of complex networks representing text, and from intermittency concepts where text is treated as a time series. Comparative experiments were performed with the New Testament in 15 different languages and with distinct books in English and Portuguese in order to quantify the dependency of the different measurements on the language and on the story being told in the book. The metrics found to be informative in distinguishing real texts from their shuffled versions include assortativity, degree and selectivity of words. As an illustration, we analyze an undeciphered medieval manuscript known as the Voynich Manuscript. We show that it is mostly compatible with natural languages and incompatible with random texts. We also obtain candidates for key-words of the Voynich Manuscript which could be helpful in the effort of deciphering it. Because we were able to identify statistical measurements that are more dependent on the syntax than on the semantics, the framework may also serve for text analysis in language-dependent applications.
References in corpus (10)
- Power-law distributions in empirical data
- Beyond word frequency: Bursts, lulls, and scaling in the temporal distributions of words
- Languages cool as they expand: Allometric scaling and the decreasing need for new words
- Network properties of written human language
- On the origin of long-range correlations in texts
- Statistical keyword detection in literary corpora
- Identification of Literary Movements Using Complex Networks to Represent Texts
- Beyond the average: Detecting global singular nodes from local features in complex networks
- Automatic Network Fingerprinting through Single-Node Motifs
- Using Complex Networks to Quantify Consistency in the Use of Words
Cited by in corpus (26)
- Text authorship identified using the dynamics of word co-occurrence networks
- A complex network approach to stylometry
- Word sense disambiguation: a complex network approach
- Probing the topological properties of complex networks modeling short written texts
- Statistical laws in linguistics
- Classifying informative and imaginative prose using complex networks
- Comparing the writing style of real and artificial papers
- Using word embeddings to improve the discriminability of co-occurrence text networks
- Concentric network symmetry grasps authors' styles in word adjacency networks
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Authorship Attribution Based on Life-Like Network Automata
- Paragraph-based complex networks: application to document classification and authenticity verification
- Representation of texts as complex networks: a mesoscopic approach
- Topic segmentation via community detection in complex networks
- Authorship attribution via network motifs identification
- Complexity-entropy analysis at different levels of organization in written language
- Network analysis of named entity co-occurrences in written texts
- Labelled network subgraphs reveal stylistic subtleties in written texts
- Semantic flow in language networks
- An Image Analysis Approach to the Calligraphy of Books
- Extractive Multi Document Summarization using Dynamical Measurements of Complex Networks
- A perspective on the advancement of natural language processing tasks via topological analysis of complex networks
- Topic Modeling in the Voynich Manuscript
- Character Entropy in Modern and Historical Texts: Comparison Metrics for an Undeciphered Manuscript
- Language Networks: a Practical Approach
- Universal and non-universal text statistics: Clustering coefficient for language identification