Comparing intermittency and network measurements of words and their dependency on authorship
arXiv:1112.6045 · doi:10.1088/1367-2630/13/12/123024
Abstract
Many features from texts and languages can now be inferred from statistical analyses using concepts from complex networks and dynamical systems. In this paper we quantify how topological properties of word co-occurrence networks and intermittency (or burstiness) in word distribution depend on the style of authors. Our database contains 40 books from 8 authors who lived in the 19th and 20th centuries, for which the following network measurements were obtained: clustering coefficient, average shortest path lengths, and betweenness. We found that the two factors with stronger dependency on the authors were the skewness in the distribution of word intermittency and the average shortest paths. Other factors such as the betweeness and the Zipf's law exponent show only weak dependency on authorship. Also assessed was the contribution from each measurement to authorship recognition using three machine learning methods. The best performance was a ca. 65 % accuracy upon combining complex network and intermittency features with the nearest neighbor algorithm. From a detailed analysis of the interdependence of the various metrics it is concluded that the methods used here are complementary for providing short- and long-scale perspectives of texts, which are useful for applications such as identification of topical words and information retrieval.
References in corpus (9)
- Characterization of complex networks: A survey of measurements
- Characterization and Modeling of weighted networks
- Beyond word frequency: Bursts, lulls, and scaling in the temporal distributions of words
- Parameter estimation for power-law distributions by maximum likelihood methods
- Network properties of written human language
- Strong correlations between text quality and complex networks features
- Statistical keyword detection in literary corpora
- Correlations between structure and dynamics in complex networks
- The meta book and size-dependent properties of written language
Cited by in corpus (25)
- A systematic comparison of supervised classifiers
- Measuring the evolution of contemporary western popular music
- Text authorship identified using the dynamics of word co-occurrence networks
- A complex network approach to stylometry
- Probing the topological properties of complex networks modeling short written texts
- Extractive Multi-document Summarization Using Multilayer Networks
- Probing the statistical properties of unknown texts: application to the Voynich Manuscript
- Classifying informative and imaginative prose using complex networks
- Comparing the writing style of real and artificial papers
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Authorship Attribution Based on Life-Like Network Automata
- Paragraph-based complex networks: application to document classification and authenticity verification
- On the use of topological features and hierarchical characterization for disambiguating names in collaborative networks
- Identification of Literary Movements Using Complex Networks to Represent Texts
- On the role of words in the network structure of texts: application to authorship attribution
- Topological-collaborative approach for disambiguating authors' names in collaborative networks
- Complex networks analysis of language complexity
- Representation of texts as complex networks: a mesoscopic approach
- Unveiling the relationship between complex networks metrics and word senses
- Topic segmentation via community detection in complex networks
- Authorship attribution via network motifs identification
- On predicting research grants productivity
- Using Complex Networks to Quantify Consistency in the Use of Words
- Labelled network subgraphs reveal stylistic subtleties in written texts
- A perspective on the advancement of natural language processing tasks via topological analysis of complex networks