Definition
The empirical study of systematic lexical co‑occurrence patterns within a corpus—measuring the degree to which words or morphemes appear together more (or less) often than expected by chance within a defined context window or syntactic relation—to identify phraseological units, idioms, and lexical associations that inform semantics and usage.
Principle
Principle
A statistically reliable deviation from expected co‑occurrence frequency indicates an associative relationship between lexical items; the strength and interpretation of that association depend on the chosen window, statistic, and corpus scope.
Demonstration
Demonstration
Illustrative scenario: A corpus analyst computes association scores across a large news corpus and finds that the lemma pair "climate – change" has high association within a two‑word adjacency window and appears with verbs like "mitigate" and "accelerate"; the analyst interprets this as a stable phraseological unit and inspects concordances to confirm contextual usages.
Misapplication
Misapplication
Assuming adjacency equals collocation without specifying window size or significance thresholds, or treating low-frequency high‑score pairs as robust associations; the error is conflating measurement artifacts (window/statistic) with stable lexical relations.
Consequence
Consequence
Collocation analysis yields evidence for lexicographic entries, multiword expression extraction, and feature engineering in NLP; incorrect parameterization or unrepresentative corpora produce spurious associations that misinform lexicon building and downstream models.
Reversal
Reversal
In languages with free word order, rich morphology, or where semantic associations manifest across syntactic dependencies rather than surface adjacency, adjacency‑based measures can fail and dependency‑ or lemma‑based association measures become necessary.
Boundary
Boundary
Clearly within: statistically significant co‑occurrence of lexical items within defined windows or dependency relations. Boundary case: associations prominent in a subcorpus but absent in general corpora (register‑specific collocations). Clearly outside: topical co‑occurrence across documents that lack syntactic or proximal relation.
Semantic Tension
Semantic Tension
Statistical significance and frequency-based detection ↔ interpretive relevance and functional collocational role in grammar and meaning.
Synthesis
Synthesis
Collocation analysis is a method for making lexical association explicit and measurable, but its findings must be interpreted through choices of window, statistic and corpus register to connect quantitative association with functional linguistic relevance.