All papers
Phonetics & linguisticsFull-paper digest

Vowel Allophony Improves Maximum-Likelihood Classification of Warlpiri Consonants

Coralie Cram, John McGahay, Megha Sundara

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.0 KB · Ready to paste

Preview copied content

TL;DR — This paper investigates acoustic cues for distinguishing five-way stop place contrasts in Warlpiri using maximum-likelihood classification, demonstrating that consonant-intrinsic cues alone are insufficient and that adjacent vowel allophony (midpoint formants) significantly improves classification.

Key contributions

  • Evaluated a maximum-likelihood classifier framework on semi-spontaneous speech from a low-resource Australian language (Warlpiri) across 5 places of articulation.
  • Systematically tested combinations of 31 acoustic cues split into stop-intrinsic (S), formant transition (T), and vowel midpoint allophony (M) groups.
  • Demonstrated that adding vowel midpoint cues (M) to intrinsic cues (S) yields a larger improvement in consonant classification than adding formant transitions (T) alone, particularly in CV sequences.
  • Established that incorporating vowel midpoint information does not degrade vowel classification accuracy while substantially reducing consonant confusability.

Problem

Australian languages typically feature a dense place-of-articulation inventory for stops (4 to 6 places) combined with fewer manner contrasts, creating a crowded perceptual continuum. Traditional consonant-intrinsic cues and F2 transitions often fail to cleanly separate adjacent places like alveolars, retroflexes, dentals, and palatals due to acoustic overlap. This paper addresses how optimal cue integration allows native speakers to overcome these high confusability levels, providing a computational window into speech perception for low-resource languages.

Method

The study utilized semi-spontaneous monologic narratives from DoReCo Warlpiri field recordings (14 female speakers, aged 25-50), processed with MAUS forced alignment and hand-corrected boundaries. A set of 31 acoustic cues was extracted using Praat and AutoVOT, grouped into stop-intrinsic cues (S: duration, burst spectral moments COG, SD, skewness, kurtosis, E-H/M, relative intensity, and F1-F4 locus at boundaries), formant transition cues (T: F1-F4 change in the 25% vowel portion adjacent to the consonant), and vowel midpoint formants (M: F1-F4 at the vowel midpoint representing allophonic quality).

Maximum-likelihood classifiers modeled each token category (CV or VC sequences) using multivariate Gaussian distributions parameterized by sample means and covariances over all logical combinations of cue sets (S, T, M), resulting in 7 models per sequence type. Classification performance was evaluated using stratified 10-fold cross-validation with macro-averaged F-scores. The architecture leveraged probabilistic modeling to simulate optimal Bayesian listener performance, testing whether extrinsic coarticulatory and allophonic cues could resolve acoustic ambiguities present in the consonant closure and release bursts.

Experimental setup

Evaluated on 4,755 tokens for CV sequences and 4,801 tokens for VC sequences extracted from Warlpiri VCV contexts involving five stops (/p, t, ó, c, k/). The framework compared 7 distinct cue-set configurations (S, T, M, ST, SM, TM, STM) using stratified 10-fold cross-validation, reporting macro-averaged F-scores for consonants, vowels, and whole sequences, with significance assessed via permutation tests and Bonferroni correction (alpha < 0.005).

Results

Consonants were highly confusable using stop-intrinsic cues alone, yielding a baseline consonant F-score of 0.654 for CVs and 0.647 for VCs. Adding vowel midpoint information (SM model) significantly boosted CV consonant F-score to 0.702 (d = 0.048, p = 0.0011), outperforming the transition-only addition (ST model at 0.681; d = 0.021, p = 0.0015). For VCs, both midpoint and transition additions improved consonant classification to roughly 0.681 and 0.676 respectively, though VC models generally performed slightly worse than CV models overall.

Cue SetCV Consonant F-scoreCV Vowel F-scoreVC Consonant F-scoreVC Vowel F-score
S0.6540.6890.6470.587
T0.2760.5590.3560.458
M0.3120.8150.5080.769
ST0.6810.7980.6760.740
SM0.7020.8100.6810.748
STM0.6910.8120.6910.760

Limitations

The study is limited to a single under-resourced Australian language (Warlpiri) with a specific 3-vowel inventory, restricting immediate generalizability to languages with larger vowel spaces. The dataset was restricted to 14 female speakers from archival recordings, excluding male speaker variability. Automatic formant tracking and burst detection via AutoVOT and Praat introduced minor data loss due to extraction errors.

Why read this

Speech and ML researchers studying low-resource phonetic analysis or Bayesian models of speech perception should read this to see how maximum-likelihood classifiers on semi-spontaneous corpora can quantify cue utility and evaluate the role of vowel allophony in place distinctions.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Computational speech perception modeling, automated acoustic analysis for low-resource languages, and linguistic fieldwork validation.

Institutions

UCLA

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3107