Definition
The estimated probability of occurrence of a particular sequence of phonological segments in a language, derived from segmental distributions and conditional sequence frequencies in corpora or phonological grammars; used as a quantitative descriptor of wordlikeness and as an input to models of perception, lexical access, and speech error likelihood.

Principle

Principle
Phonotactic probability operationalizes the relative commonness of segment sequences: sequences with higher estimated probability are modeled as more likely to occur or be judged wordlike by systems that use distributional evidence; these estimates depend on the chosen unit (phoneme, biphone, onset) and corpus or model used to derive frequencies.

Demonstration

Demonstration
Illustrative scenario → Situation: a researcher builds a model of English onset legality for nonce‑word experiments. Recognition: compute frequencies of onset clusters in a representative phoneme corpus and estimate conditional probabilities for target clusters. Action: use these probabilities to predict which nonce clusters participants will accept as wordlike. Consequence: clusters with higher estimated probabilities are expected to receive higher wordlikeness ratings in that experimental context, all else equal.

Misapplication

Misapplication
Applying phonotactic probabilities computed in one language or register to another without adjustment, or conflating high phonotactic probability with semantic plausibility; the error lies in treating distributional likelihood as universal well‑formedness or meaning.

Consequence

Consequence
Phonotactic probability estimates inform models in speech perception, psycholinguistic experiments, language acquisition modelling and ASR/lexicon weighting; inappropriate estimates can bias predictions about acceptability, lexical competition, or recognition performance.

Reversal

Reversal
In loanword adaptation, proper names, onomatopoeia, or morphological boundary cases, distributional phonotactic estimates may not predict acceptability or occurrence; language-specific phonological rules and morphophonological structure can override surface sequence probabilities.

Boundary

Boundary
Clearly within: segmental sequence probability estimates derived from representative corpora for predicting distributional tendency and wordlikeness. Boundary case: sequences crossing morpheme boundaries where morphological composition affects probability. Clearly outside: suprasegmental properties (stress, tone) or higher‑level prosodic well‑formedness not captured by segmental sequence probability alone.

Semantic Tension

Semantic Tension
Frequency‑based statistical description (phonotactic probability) ↔ categorical phonological constraints or rule‑based grammars; statistical regularities predict tendencies, while categorical constraints may forbid sequences regardless of corpus frequency.

Synthesis

Synthesis
Phonotactic probability is a descriptive, model‑dependent measure that provides a probabilistic proxy for phonological well‑formedness and processing likelihood; it complements but does not replace categorical phonological analysis and must be interpreted relative to the derivation method and linguistic context.