 ##  [Phonotactic Probability](/phonotactic-probability-0) 

 Definition

The estimated probability of occurrence of a particular sequence of phonological segments in a language, derived from segmental distributions and conditional sequence frequencies in corpora or phonological grammars; used as a quantitative descriptor of wordlikeness and as an input to models of perception, lexical access, and speech error likelihood.

 

 

 

 

 

 





## Principle

Principle

Phonotactic probability operationalizes the relative commonness of segment sequences: sequences with higher estimated probability are modeled as more likely to occur or be judged wordlike by systems that use distributional evidence; these estimates depend on the chosen unit (phoneme, biphone, onset) and corpus or model used to derive frequencies.

 

 

 

 

 





## Demonstration

Demonstration

Illustrative scenario → Situation: a researcher builds a model of English onset legality for nonce‑word experiments. Recognition: compute frequencies of onset clusters in a representative phoneme corpus and estimate conditional probabilities for target clusters. Action: use these probabilities to predict which nonce clusters participants will accept as wordlike. Consequence: clusters with higher estimated probabilities are expected to receive higher wordlikeness ratings in that experimental context, all else equal.

 

 

 

 

## Misapplication

Misapplication

Applying phonotactic probabilities computed in one language or register to another without adjustment, or conflating high phonotactic probability with semantic plausibility; the error lies in treating distributional likelihood as universal well‑formedness or meaning.

 

 

 

 

 





## Consequence

Consequence

Phonotactic probability estimates inform models in speech perception, psycholinguistic experiments, language acquisition modelling and ASR/lexicon weighting; inappropriate estimates can bias predictions about acceptability, lexical competition, or recognition performance.

 

 

 

 

## Reversal

Reversal

In loanword adaptation, proper names, onomatopoeia, or morphological boundary cases, distributional phonotactic estimates may not predict acceptability or occurrence; language-specific phonological rules and morphophonological structure can override surface sequence probabilities.

 

 

 

 

 





## Boundary

Boundary

Clearly within: segmental sequence probability estimates derived from representative corpora for predicting distributional tendency and wordlikeness. Boundary case: sequences crossing morpheme boundaries where morphological composition affects probability. Clearly outside: suprasegmental properties (stress, tone) or higher‑level prosodic well‑formedness not captured by segmental sequence probability alone.

 

 

 

 

 





## Semantic Tension

Semantic Tension

Frequency‑based statistical description (phonotactic probability) ↔ categorical phonological constraints or rule‑based grammars; statistical regularities predict tendencies, while categorical constraints may forbid sequences regardless of corpus frequency.

 

 

 

 

 





## Synthesis

Synthesis

Phonotactic probability is a descriptive, model‑dependent measure that provides a probabilistic proxy for phonological well‑formedness and processing likelihood; it complements but does not replace categorical phonological analysis and must be interpreted relative to the derivation method and linguistic context.