Definition
The mean linear distance, measured in number of intervening word tokens, between syntactically linked head and dependent tokens across a dependency-parsed dataset; computed as the average of absolute differences in token positions for all annotated dependencies and used as an index of syntactic locality.
Principle
Principle
Shorter mean dependency distance indicates greater linear locality of syntactic relations, which correlates with lower short-term memory load and faster incremental processing in many experimental and corpus studies; longer distances reflect greater separation between dependent elements and potential increased processing demands.
Demonstration
Demonstration
Illustrative scenario → Sentence 1 (adjacent subject–verb): “Cats sleep.” Subject and verb distance = 1. Sentence 2 (intervening relative clause): “The cat that the dog chased sleeps.” Subject–verb distance larger because of intervening material; the mean dependency distance computed across dependencies will be higher for sentences like Sentence 2, predicting greater processing cost under locality assumptions.
Misapplication
Misapplication
Treating a lower mean dependency distance as proof of syntactic simplicity without considering dependency density, multiple embedding, annotation scheme differences, or the role of morphological marking and prosodic cues that can offset linear separation.
Consequence
Consequence
Dependency distance is used in typological comparisons, register studies and psycholinguistic predictions; aggregated increases in mean dependency distance across corpora suggest greater average linear separation of arguments and may correlate with longer reading times or higher cognitive load in processing models, subject to annotation and genre controls.
Reversal
Reversal
In languages with free word order, extensive case marking, rich morphology, or strong prosodic cues, linear distance may be a poorer predictor of processing cost; non-projective dependencies or differing annotation conventions also complicate the interpretation of linear-distance measures.
Boundary
Boundary
Clearly within: corpora annotated with a consistent dependency scheme and tokenization where distance is measured as absolute token-position difference. Boundary case: cross-corpus comparisons where parsers or tokenization differ. Clearly outside: constituency-based measures of syntactic depth that do not use linear head–dependent distance.
Semantic Tension
Semantic Tension
Dependency Distance ↔ Information Density — minimizing linear distance reduces locality cost but may concentrate informational content per word, creating a trade-off between locality and per-token information load.
Synthesis
Synthesis
Mean dependency distance is a reproducible, quantifiable proxy for syntactic locality and an informative predictor in many contexts, but claims about cognitive processing require controlled annotation, consideration of morphological and prosodic cues, and complementary evidence.