Dependency-distance estimates derived from a single corpus are often considered properties of a language. However, analysis across 38 same-language treebank pairs shows only moderate agreement, with substitution reversing nearly 40% of language orderings.
The study used concordance correlation, Bland-Altman analysis, and a multiverse design, revealing that treebank choice accounts for about 29% of variance. This disagreement exceeds within-treebank sampling error and persists across multiple preprocessing specifications.
Despite the variability, all treebanks confirmed dependency-length minimization, indicating some universal aspect. The findings suggest that mean dependency-distance is more a corpus-conditioned composite influenced by grammatical, register, and annotation factors than a stable language-level parameter.
This impacts how dependency-distance metrics are used in linguistic and computational models, emphasizing the importance of corpus selection and preprocessing choices.
Source: https://arxiv.org/abs/2609.04223