Definition
The estimated probability of occurrence of a particular sequence of phonological segments in a language, derived from segmental distributions and conditional sequence frequencies in corpora or phonological grammars; used as a quantitative descriptor of wordlikeness and as an input to models of perception, lexical access, and speech error likelihood.
Principle
Principle
Phonotactic probability operationalizes the relative commonness of segment sequences: sequences with higher estimated probability are modeled as more likely to occur or be judged wordlike by systems that use distributional evidence; these estimates depend on the chosen unit (phoneme, biphone, onset) and corpus or model used to derive frequencies.
Demonstration
Demonstration
Illustrative scenario → Situation: a researcher builds a model of English onset legality for nonce‑word experiments. Recognition: compute frequencies of onset clusters in a representative phoneme corpus and estimate conditional probabilities for target clusters. Action: use these probabilities to predict which nonce clusters participants will accept as wordlike. Consequence: clusters with higher estimated probabilities are expected to receive higher wordlikeness ratings in that experimental context, all else equal.
Misapplication
Misapplication
Applying phonotactic probabilities computed in one language or register to another without adjustment, or conflating high phonotactic probability with semantic plausibility; the error lies in treating distributional likelihood as universal well‑formedness or meaning.
Consequence
Consequence
Phonotactic probability estimates inform models in speech perception, psycholinguistic experiments, language acquisition modelling and ASR/lexicon weighting; inappropriate estimates can bias predictions about acceptability, lexical competition, or recognition performance.
Reversal
Reversal
In loanword adaptation, proper names, onomatopoeia, or morphological boundary cases, distributional phonotactic estimates may not predict acceptability or occurrence; language-specific phonological rules and morphophonological structure can override surface sequence probabilities.
Boundary
Boundary
Clearly within: segmental sequence probability estimates derived from representative corpora for predicting distributional tendency and wordlikeness. Boundary case: sequences crossing morpheme boundaries where morphological composition affects probability. Clearly outside: suprasegmental properties (stress, tone) or higher‑level prosodic well‑formedness not captured by segmental sequence probability alone.
Semantic Tension
Semantic Tension
Frequency‑based statistical description (phonotactic probability) ↔ categorical phonological constraints or rule‑based grammars; statistical regularities predict tendencies, while categorical constraints may forbid sequences regardless of corpus frequency.
Synthesis
Synthesis
Phonotactic probability is a descriptive, model‑dependent measure that provides a probabilistic proxy for phonological well‑formedness and processing likelihood; it complements but does not replace categorical phonological analysis and must be interpreted relative to the derivation method and linguistic context.