Skip to content

LLMs1 min read

Corpus choice influences dependency-distance estimates across languages

Dependency-distance estimates vary significantly across corpora, with nearly 40% of language orderings reversed when substituting treebanks. This suggests corpus factors impact these measurements more than language properties.

By OpenSmartRoute editorial · written through the router by llm-onprem

From arXiv cs.CL - “How Much Does Corpus Choice Change Dependency-Distance Estimates?

Dependency-distance estimates derived from a single corpus are often considered properties of a language. However, analysis across 38 same-language treebank pairs shows only moderate agreement, with substitution reversing nearly 40% of language orderings.

The study used concordance correlation, Bland-Altman analysis, and a multiverse design, revealing that treebank choice accounts for about 29% of variance. This disagreement exceeds within-treebank sampling error and persists across multiple preprocessing specifications.

Despite the variability, all treebanks confirmed dependency-length minimization, indicating some universal aspect. The findings suggest that mean dependency-distance is more a corpus-conditioned composite influenced by grammatical, register, and annotation factors than a stable language-level parameter.

This impacts how dependency-distance metrics are used in linguistic and computational models, emphasizing the importance of corpus selection and preprocessing choices.

Source: https://arxiv.org/abs/2609.04223

Published Sep 7, 2026 · updated Sep 7, 2026 · 124 words

Keep reading

Related posts

More in LLMs