Topic segmentation via community detection in complex networks
arXiv:1512.01384 · doi:10.1063/1.4954215
Abstract
Many real systems have been modelled in terms of network concepts, and written texts are a particular example of information networks. In recent years, the use of network methods to analyze language has allowed the discovery of several interesting findings, including the proposition of novel models to explain the emergence of fundamental universal patterns. While syntactical networks, one of the most prevalent networked models of written texts, display both scale-free and small-world properties, such representation fails in capturing other textual features, such as the organization in topics or subjects. In this context, we propose a novel network representation whose main purpose is to capture the semantical relationships of words in a simple way. To do so, we link all words co-occurring in the same semantic context, which is defined in a threefold way. We show that the proposed representations favours the emergence of communities of semantically related words, and this feature may be used to identify relevant topics. The proposed methodology to detect topics was applied to segment selected Wikipedia articles. We have found that, in general, our methods outperform traditional bag-of-words representations, which suggests that a high-level textual representation may be useful to study semantical features of texts.
References in corpus (10)
- Fast unfolding of communities in large networks
- Uncovering the overlapping community structure of complex networks in nature and society
- A complex network approach to stylometry
- Statistical keyword detection in literary corpora
- Comparing the writing style of real and artificial papers
- Wikipedia information flow analysis reveals the scale-free architecture of the Semantic Space
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Three-feature model to reproduce the topology of citation networks and the effects from authors' visibility on their h-index
- Complex networks analysis of language complexity
- Information-theoretical analysis of the statistical dependencies among three variables: Applications to written language
Cited by in corpus (5)
- Paragraph-based complex networks: application to document classification and authenticity verification
- Representation of texts as complex networks: a mesoscopic approach
- Multilayer Networks for Text Analysis with Multiple Data Types
- The Dynamics of Knowledge Acquisition via Self-Learning in Complex Networks
- Term-community-based topic detection with variable resolution