Definition
Quantitative analysis of measurable linguistic features (e.g., function‑word frequencies, character n‑grams, sentence length distributions) using statistical or machine‑learning methods to characterise stylistic patterns for purposes such as authorship attribution, stylistic clustering, or diachronic change detection.

Principle

Principle
Authors and genres tend to produce reproducible distributions of low‑level linguistic features; by extracting and modelling these features across representative samples, analysts can estimate probabilistic affinities or attributions, conditional on sample size, genre, and model validity.

Demonstration

Demonstration
Illustrative scenario → Situation: An anonymous essay is suspected to match one of three candidate authors. Recognition: Analysts extract function‑word frequencies and character n‑grams from the candidates’ corpora and the anonymous text. Action: A classification model computes attribution probabilities and reports uncertainty intervals. Consequence: Stylometric evidence contributes a probabilistic attribution that must be reported with caveats and evaluated alongside other evidence.

Misapplication

Misapplication
Treating stylometric output as definitive proof — for example, using a model trained on small or topically homogeneous samples, ignoring genre/topic confounds, or failing to report uncertainty — overstates evidential value and risks false attribution.

Consequence

Consequence
Properly used, stylometry provides replicable, quantitative evidence about textual affinities that can strengthen multidisciplinary attribution arguments; misused, it produces overconfident, potentially incorrect attributions and obscures alternative explanations (editing, collaboration, translation).

Reversal

Reversal
Signals weaken or become unreliable for very short texts, highly edited or collaborative works, translations, or when authors deliberately obfuscate style; in such cases stylometric inference may fail or require different feature sets and robust uncertainty quantification.

Boundary

Boundary
Clearly within: comparative analysis of multi‑thousand‑word texts where feature distributions are stable. Boundary case: mid‑length texts (few hundred words) where feature variance is large. Clearly outside: close interpretive reading focused on thematic or symbolic meaning rather than measurable surface features.

Semantic Tension

Semantic Tension
Tension with Close Reading — stylometry privileges measurable, often low‑level features and reproducibility, while close reading privileges interpretive nuance and contextual meaning; both approaches address different questions and can complement each other when integrated carefully.

Synthesis

Synthesis
Stylometry is a probabilistic, empirical tool: it quantifies stylistic signals that can inform authorship and stylistic questions but must be applied with sufficient data, explicit modelling choices, and transparent uncertainty to avoid overclaiming.