Text authorship identified using the dynamics of word co-occurrence networks
arXiv:1608.01965 · doi:10.1371/journal.pone.0170527
Abstract
The identification of authorship in disputed documents still requires human expertise, which is now unfeasible for many tasks owing to the large volumes of text and authors in practical applications. In this study, we introduce a methodology based on the dynamics of word co-occurrence networks representing written texts to classify a corpus of 80 texts by 8 authors. The texts were divided into sections with equal number of linguistic tokens, from which time series were created for 12 topological metrics. The series were proven to be stationary (p-value>0.05), which permits to use distribution moments as learning attributes. With an optimized supervised learning procedure using a Radial Basis Function Network, 68 out of 80 texts were correctly classified, i.e. a remarkable 85% author matching success rate. Therefore, fluctuations in purely dynamic network metrics were found to characterize authorship, thus opening the way for the description of texts in terms of small evolving networks. Moreover, the approach introduced allows for comparison of texts with diverse characteristics in a simple, fast fashion.
References in corpus (14)
- Criticality of spreading dynamics in hierarchical cluster networks without inhibition
- Probing the topological properties of complex networks modeling short written texts
- Wikipedia information flow analysis reveals the scale-free architecture of the Semantic Space
- Word sense disambiguation via high order of learning in complex networks
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Three-feature model to reproduce the topology of citation networks and the effects from authors' visibility on their h-index
- Identification of Literary Movements Using Complex Networks to Represent Texts
- On time-varying collaboration networks
- Complex networks analysis of language complexity
- Extracting information from S-curves of language change
- Modeling the average shortest path length in growth of word-adjacency networks
- A stronger null hypothesis for crossing dependencies
- Extracting directed information flow networks: an application to genetics and semantics
- Modeling the Co-occurrence Principles of the Consonant Inventories: A Complex Network Approach