The risks of mixing dependency lengths from sequences of different length
arXiv:1304.3841 · doi:10.1515/glot-2014-0014
Abstract
Mixing dependency lengths from sequences of different length is a common practice in language research. However, the empirical distribution of dependency lengths of sentences of the same length differs from that of sentences of varying length and the distribution of dependency lengths depends on sentence length for real sentences and also under the null hypothesis that dependencies connect vertices located in random positions of the sequence. This suggests that certain results, such as the distribution of syntactic dependency lengths mixing dependencies from sentences of varying length, could be a mere consequence of that mixing. Furthermore, differences in the global averages of dependency length (mixing lengths from sentences of varying length) for two different languages do not simply imply a priori that one language optimizes dependency lengths better than the other because those differences could be due to differences in the distribution of sentence lengths and other factors.
Laguage and referencing has been improved; Eqs. 7, 11, B7 and B8 have been corrected
Cited by in corpus (16)
- Contrasting Linguistic Patterns in Human and LLM-Generated News Text
- The placement of the head that minimizes online memory: a complex systems approach
- Are crossing dependencies really scarce?
- The scarcity of crossing dependencies: a direct outcome of a specific constraint?
- Anti dependency distance minimization in short sequences. A graph theoretic approach
- Non-crossing dependencies: least effort, not grammar
- Memory limitations are hidden in grammar
- Crossings as a side effect of dependency lengths
- The scaling of the minimum sum of edge lengths in uniformly random trees
- Towards a theory of word order. Comment on "Dependency distance: a new perspective on syntactic patterns in natural language" by Haitao Liu et al
- Dependency length minimization: Puzzles and Promises
- Bounds of the sum of edge lengths in linear arrangements of trees
- Beyond description. Comment on "Approaching human language with complex networks" by Cong & Liu
- Inherent Dependency Displacement Bias of Transition-Based Algorithms
- A commentary on "The now-or-never bottleneck: a fundamental constraint on language", by Christiansen and Chater (2016)
- The Impact of Edge Displacement Vaserstein Distance on UD Parsing Performance