Definition
A quantitative measure expressing the proportion of content-bearing lexical items (typically nouns, main verbs, adjectives and adverbs) to the total number of words or tokens in a stretch of text, used as an index of informational concentration per word.
Principle
Principle
Texts with higher proportions of content words tend to package more informational content per token, but the metric depends on how words and categories are segmented and defined in a given language or annotation scheme.
Demonstration
Demonstration
Illustrative scenario → Situation: Two short paragraphs are compared; Recognition: analyst tags tokens as content or function words; Action: compute content-word count ÷ total tokens for each paragraph; Consequence: the paragraph with the higher ratio is described as having greater lexical density, suggesting higher information per word under that coding scheme.
Misapplication
Misapplication
Interpreting lexical density as a direct proxy for readability or stylistic quality without accounting for genre, purpose, or morphological structure; this confuses informational concentration with comprehensibility or aesthetic value.
Consequence
Consequence
Lexical density is useful for comparing registers, genres or translation choice, and for corpus profiling; it can mislead when applied across languages with different morphological typologies or when tokenization/tagging rules vary.
Reversal
Reversal
In agglutinative or polysynthetic languages, information is often carried in bound morphemes rather than separate word tokens, so word-based lexical density may underrepresent informational concentration; conversely, analytic languages may show different distributions.
Boundary
Boundary
Clearly within: a prose paragraph analyzed with an explicit part-of-speech tagset counting content-word tokens; Boundary case: texts with many contractions or clitics where tokenization choices affect counts; Clearly outside: qualitative claims about 'depth' or 'complexity' that ignore measurable token proportions.
Semantic Tension
Semantic Tension
Information Density ↔ Readability — higher lexical density indicates more information per token but can conflict with ease of processing for particular readers or communicative aims.
Synthesis
Synthesis
Lexical density is a descriptive, contingent metric: it quantifies one aspect of textual informational structure under a specific tokenization and tagging regime, and so should be interpreted relative to language, genre and method.