In narrative texts punctuation marks obey the same statistics as words
arXiv:1604.00834 · doi:10.1016/j.ins.2016.09.051
Abstract
From a grammar point of view, the role of punctuation marks in a sentence is formally defined and well understood. In semantic analysis punctuation plays also a crucial role as a method of avoiding ambiguity of the meaning. A different situation can be observed in the statistical analyses of language samples, where the decision on whether the punctuation marks should be considered or should be neglected is seen rather as arbitrary and at present it belongs to a researcher's preference. An objective of this work is to shed some light onto this problem by providing us with an answer to the question whether the punctuation marks may be treated as ordinary words and whether they should be included in any analysis of the word co-occurences. We already know from our previous study (S.~Drożdż {\it et al.}, Inf. Sci. 331 (2016) 32-44) that full stops that determine the length of sentences are the main carrier of long-range correlations. Now we extend that study and analyze statistical properties of the most common punctuation marks in a few Indo-European languages, investigate their frequencies, and locate them accordingly in the Zipf rank-frequency plots as well as study their role in the word-adjacency networks. We show that, from a statistical viewpoint, the punctuation marks reveal properties that are qualitatively similar to the properties of the most frequent words like articles, conjunctions, pronouns, and prepositions. This refers to both the Zipfian analysis and the network analysis. By adding the punctuation marks to the Zipf plots, we also show that these plots that are normally described by the Zipf-Mandelbrot distribution largely restore the power-law Zipfian behaviour for the most frequent items.
Information Sciences (inprint)
References in corpus (5)
- Beyond word frequency: Bursts, lulls, and scaling in the temporal distributions of words
- Network properties of written human language
- Structure-semantics interplay in complex networks and its effects on the predictability of similarity in texts
- Modeling the average shortest path length in growth of word-adjacency networks
- Network model of human language
Cited by in corpus (15)
- Quantitative approach to multifractality induced by correlations and broad distribution of data
- Using word embeddings to improve the discriminability of co-occurrence text networks
- Complex systems approach to natural language
- Linguistic data mining with complex networks: a stylometric-oriented approach
- Hierarchical organization of H. Eugene Stanley scientific collaboration community in weighted network representation
- Universal versus system-specific features of punctuation usage patterns in~major Western~languages
- Text characterization based on recurrence networks
- Statistics of punctuation in experimental literature -- the remarkable case of "Finnegans Wake" by James Joyce
- Entropy and type-token ratio in gigaword corpora
- Multifractal hopscotch in "Hopscotch" by Julio Cortazar
- Punctuation patterns in "Finnegans Wake" by James Joyce are largely translation-invariant
- Quantifying patterns of punctuation in modern Chinese prose
- Average shortest-path length in word-adjacency networks: Chinese versus English
- A Zipf-preserving, long-range correlated surrogate for written language and other symbolic sequences
- Quadratic Term Correction on Heaps' Law