Classifying informative and imaginative prose using complex networks
arXiv:1507.07826 · doi:10.1209/0295-5075/113/28007
Abstract
Statistical methods have been widely employed in recent years to grasp many language properties. The application of such techniques have allowed an improvement of several linguistic applications, which encompasses machine translation, automatic summarization and document classification. In the latter, many approaches have emphasized the semantical content of texts, as it is the case of bag-of-word language models. This approach has certainly yielded reasonable performance. However, some potential features such as the structural organization of texts have been used only on a few studies. In this context, we probe how features derived from textual structure analysis can be effectively employed in a classification task. More specifically, we performed a supervised classification aiming at discriminating informative from imaginative documents. Using a networked model that describes the local topological/dynamical properties of function words, we achieved an accuracy rate of up to 95%, which is much higher than similar networked approaches. A systematic analysis of feature relevance revealed that symmetry and accessibility measurements are among the most prominent network measurements. Our results suggest that these measurements could be used in related language applications, as they play a complementary role in characterizing texts.
References in corpus (9)
- A systematic comparison of supervised classifiers
- Probing the statistical properties of unknown texts: application to the Voynich Manuscript
- Word sense disambiguation via high order of learning in complex networks
- Concentric network symmetry grasps authors' styles in word adjacency networks
- Detecting degree symmetries in networks
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Identification of Literary Movements Using Complex Networks to Represent Texts
- Complex networks analysis of language complexity
- Unveiling the relationship between complex networks metrics and word senses
Cited by in corpus (15)
- Using word embeddings to improve the discriminability of co-occurrence text networks
- Paragraph-based complex networks: application to document classification and authenticity verification
- On the role of words in the network structure of texts: application to authorship attribution
- A Neural Entity Coreference Resolution Review
- Representation of texts as complex networks: a mesoscopic approach
- Identifying significant edges via neighborhood information
- A complex network approach to political analysis: application to the Brazilian Chamber of Deputies
- A comparative analysis of knowledge acquisition performance in complex networks
- Analyzing the relationship between text features and research proposal productivity
- Semantic flow in language networks
- Labelled network subgraphs reveal stylistic subtleties in written texts
- Derivative of a hypergraph as a tool for linguistic pattern analysis
- Comparing the impact of subfields in scientific journals
- Language Networks: a Practical Approach
- Associations between author-level metrics in subsequent time periods