Authorship recognition via fluctuation analysis of network topology and word intermittency
arXiv:1502.01245 · doi:10.1088/1742-5468/2015/03/P03005
Abstract
Statistical methods have been widely employed in many practical natural language processing applications. More specifically, complex networks concepts and methods from dynamical systems theory have been successfully applied to recognize stylistic patterns in written texts. Despite the large amount of studies devoted to represent texts with physical models, only a few studies have assessed the relevance of attributes derived from the analysis of stylistic fluctuations. Because fluctuations represent a pivotal factor for characterizing a myriad of real systems, this study focused on the analysis of the properties of stylistic fluctuations in texts via topological analysis of complex networks and intermittency measurements. The results showed that different authors display distinct fluctuation patterns. In particular, it was found that it is possible to identify the authorship of books using the intermittency of specific words. Taken together, the results described here suggest that the patterns found in stylistic fluctuations could be used to analyze other related complex systems. Furthermore, the discovery of novel patterns related to textual stylistic fluctuations indicates that these patterns could be useful to improve the state of the art of many stylistic-based natural language processing tasks.
References in corpus (9)
- Fluctuation scaling in complex systems: Taylor's law and beyond
- Generalized Hurst exponent and multifractal function of original and translated texts mapped into frequency and length time series
- Statistical keyword detection in literary corpora
- Structure-semantics interplay in complex networks and its effects on the predictability of similarity in texts
- Word sense disambiguation via high order of learning in complex networks
- Identification of Literary Movements Using Complex Networks to Represent Texts
- Complex networks analysis of language complexity
- Unveiling the relationship between complex networks metrics and word senses
- Beyond the average: Detecting global singular nodes from local features in complex networks