Probing the topological properties of complex networks modeling short written texts
arXiv:1412.8504 · doi:10.1371/journal.pone.0118394
Abstract
In recent years, graph theory has been widely employed to probe several language properties. More specifically, the so-called word adjacency model has been proven useful for tackling several practical problems, especially those relying on textual stylistic analysis. The most common approach to treat texts as networks has simply considered either large pieces of texts or entire books. This approach has certainly worked well -- many informative discoveries have been made this way -- but it raises an uncomfortable question: could there be important topological patterns in small pieces of texts? To address this problem, the topological properties of subtexts sampled from entire books was probed. Statistical analyzes performed on a dataset comprising 50 novels revealed that most of the traditional topological measurements are stable for short subtexts. When the performance of the authorship recognition task was analyzed, it was found that a proper sampling yields a discriminability similar to the one found with full texts. Surprisingly, the support vector machine classification based on the characterization of short texts outperformed the one performed with entire books. These findings suggest that a local topological analysis of large documents might improve its global characterization. Most importantly, it was verified, as a proof of principle, that short texts can be analyzed with the methods and concepts of complex networks. As a consequence, the techniques described here can be extended in a straightforward fashion to analyze texts as time-varying complex networks.
References in corpus (9)
- Network properties of written human language
- Characterizing the network topology of the energy landscapes of atomic clusters
- Statistical keyword detection in literary corpora
- Wikipedia information flow analysis reveals the scale-free architecture of the Semantic Space
- Structure-semantics interplay in complex networks and its effects on the predictability of similarity in texts
- Word sense disambiguation via high order of learning in complex networks
- Authorship recognition via fluctuation analysis of network topology and word intermittency
- Identification of Literary Movements Using Complex Networks to Represent Texts
- Complex networks analysis of language complexity
Cited by in corpus (7)
- Text authorship identified using the dynamics of word co-occurrence networks
- Concentric network symmetry grasps authors' styles in word adjacency networks
- Writing about COVID-19 vaccines: Emotional profiling unravels how mainstream and alternative press framed AstraZeneca, Pfizer and vaccination campaigns
- Complexity-entropy analysis at different levels of organization in written language
- Cognitive networks identify the content of English and Italian popular posts about COVID-19 vaccines: Anticipation, logistics, conspiracy and loss of trust
- #lockdown: network-enhanced emotional profiling at the times of COVID-19
- God (), the first small world network