A complex network approach to stylometry
arXiv:1506.09107 · doi:10.1371/journal.pone.0136076
Abstract
Statistical methods have been widely employed to study the fundamental properties of language. In recent years, methods from complex and dynamical systems proved useful to create several language models. Despite the large amount of studies devoted to represent texts with physical models, only a limited number of studies have shown how the properties of the underlying physical systems can be employed to improve the performance of natural language processing tasks. In this paper, I address this problem by devising complex networks methods that are able to improve the performance of current statistical methods. Using a fuzzy classification strategy, I show that the topological properties extracted from texts complement the traditional textual description. In several cases, the performance obtained with hybrid approaches outperformed the results obtained when only traditional or networked methods were used. Because the proposed model is generic, the framework devised here could be straightforwardly used to study similar textual applications where the topology plays a pivotal role in the description of the interacting agents.
PLoS ONE, 2015 (to appear)
References in corpus (9)
- Network properties of written human language
- Probing the topological properties of complex networks modeling short written texts
- Statistical keyword detection in literary corpora
- Word sense disambiguation via high order of learning in complex networks
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Identification of Literary Movements Using Complex Networks to Represent Texts
- Complex networks analysis of language complexity
- Modeling the average shortest path length in growth of word-adjacency networks
- Unveiling the relationship between complex networks metrics and word senses
Cited by in corpus (26)
- Using network science and text analytics to produce surveys in a scientific topic
- Text authorship identified using the dynamics of word co-occurrence networks
- Word sense disambiguation: a complex network approach
- Probing the topological properties of complex networks modeling short written texts
- Extractive Multi-document Summarization Using Multilayer Networks
- Text-mining forma mentis networks reconstruct public perception of the STEM gender gap in social media
- Using word embeddings to improve the discriminability of co-occurrence text networks
- Authorship Attribution Based on Life-Like Network Automata
- Paragraph-based complex networks: application to document classification and authenticity verification
- Word sense induction using word embeddings and community detection in complex networks
- In narrative texts punctuation marks obey the same statistics as words
- Representation of texts as complex networks: a mesoscopic approach
- Topic segmentation via community detection in complex networks
- Authorship attribution via network motifs identification
- Rank dynamics of word usage at multiple scales
- The Dynamics of Knowledge Acquisition via Self-Learning in Complex Networks
- Complexity-entropy analysis at different levels of organization in written language
- Network analysis of named entity co-occurrences in written texts
- Forma mentis networks reconstruct how Italian high schoolers and international STEM experts perceive teachers, students, scientists, and school
- Labelled network subgraphs reveal stylistic subtleties in written texts
- Predicting language diversity with complex network
- Text characterization based on recurrence networks
- A pattern recognition approach for distinguishing between prose and poetry
- An Image Analysis Approach to the Calligraphy of Books
- Language Networks: a Practical Approach
- A Linear-complexity Multi-biometric Forensic Document Analysis System, by Fusing the Stylome and Signature Modalities