All papers
Phonetics & linguisticsFull-paper digest

Acoustic Pharyngometry as an Auditable Anchor for Cross-Speaker EMA Normalization

Daniel Friedrichs, Valeriia Vyshnevetska

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.4 KB · Ready to paste

Preview copied content

TL;DR — The paper introduces an auditable, pharyngometry-anchored normalization method for electromagnetic articulography (EMA) that reduces cross-speaker trajectory dispersion by 16.5% but fails to improve leave-one-speaker-out formant prediction.

Key contributions

  • A palate-referenced EMA normalization pipeline combining uniform palate-length scaling with a pharyngometry-derived oral landmark anchoring a low-parameter anterior-posterior (A-P) warp.
  • Extraction of robust hard-palate envelopes via A-P binning and envelope fitting to eliminate backtracking/loop artifacts.
  • Demonstration that uniform scaling accounts for the vast majority (15.7% out of 16.5%) of between-speaker trajectory dispersion reduction in DDK sequences.
  • Comprehensive evaluation showing that spatial alignment does not translate to improved speaker-independent articulatory-to-acoustic mapping.

Problem

Cross-speaker kinematic analyses using electromagnetic articulography (EMA) are severely confounded by differences in vocal-tract morphology. Standard head corrections and palate-based coordinate definitions fail to eliminate morphology-driven positional offsets across speakers. Existing alternatives range from rigid global transforms to highly flexible nonrigid mappings like thin-plate splines, but these often lack interpretability and auditability. Consequently, researchers need normalization procedures that explicitly correct for specific anatomical factors while remaining phonetically transparent and inspectable.

Method

The method processes synchronous EMA (1250 Hz) and audio (48 kHz) data from 29 German-speaking adults (14 with valid EMA palate traces and pharyngometry landmarks). EMA trajectories are rigidly head-corrected and expressed in a 2D midsagittal palate-referenced frame tracking three tongue sensors: tongue tip (TT), tongue blade (TB), and tongue dorsum (TD). A robust hard-palate envelope is estimated by binning points along the anterior-posterior (A-P) axis, extracting an upper envelope, and fitting a smooth curve. Its A-P extent defines a palate-length proxy Lpal(s)L_{pal}^{(s)} used for global scaling.

Acoustic pharyngometry provides distance-indexed area functions As(d)A_s(d) averaged across four trials. After smoothing with a Savitzky-Golay filter (polynomial order 3, window length 11), two landmarks are extracted: an oral-cavity expansion peak dpeak(s)d_{peak}^{(s)} (0-8 cm) and the oral-pharyngeal junction (OPJ) minimum dOPJ(s)d_{OPJ}^{(s)} (6-14 cm). The oral landmark is normalized as a unitless proportion upeak=(dpeak−d0)/(dOPJ−d0)u_{peak} = (d_{peak} - d_0) / (d_{OPJ} - d_0) to anchor a monotonic A-P piecewise warp in palate-normalized coordinate space u=(x−xant)/Lpalu = (x - x_{ant})/L_{pal}. The superior-inferior (S-I) axis is scaled uniformly by Lcanon/Lpal(s)L_{canon}/L_{pal}^{(s)} without warping.

Evaluations include DDK (diadochokinetic) blocks with vowels /i,a,u/ and sustained vowels (/i, y, e, ø, E, a, o, u/). Formants (F1, F2) are extracted via Burg LPC in Praat. Leave-one-speaker-out (LOSO) ridge regression (L2-regularized, tuning penalty λ\lambda via inner CV) predicts F1 and F2 from concatenated 6D tongue positions, tested across Hz, Bark, VTLN-style length scaling, and palate-relative constriction predictors. Speaker identification is tested via multinomial logistic regression across cycles, comparing full trajectories, centered positions, and velocity-only features.

Experimental setup

Evaluated on a German multimodal dataset comprising synchronous EMA and audio for 29 speakers (14 utilized after strict quality control on palate traces and pharyngometry landmarks). Data includes high-repetition DDK sequences and sustained vowels. Baselines include unnormalized baseline coordinates and scale-only uniform normalization. Metrics include mean pairwise Euclidean distance (between-speaker dispersion), Pearson correlation rr and RMSE for LOSO formant prediction, and classification accuracy for speaker identification across 5-fold cross-validation.

Results

Scale-only normalization reduced mean between-speaker dispersion from 31.28 mm to 26.36 mm (15.7% reduction; 91 speaker pairs, N=14N=14), while adding the pharyngometry-anchored A-P warp yielded a minor additional decrease to 26.13 mm (16.5% total reduction). However, LOSO ridge regression for formant prediction did not improve with the warp: F1/F2 correlations dropped compared to scale-only across all target modes (e.g., standard Hz correlation for scale-only was 0.381/0.646 versus 0.282/0.612 for scale+warp).

Speaker identification accuracy from absolute trajectory positions actually increased after normalization, rising from 0.751 (baseline) to 0.800 (scale-only) and 0.902 (scale+warp), whereas performance dropped sharply to chance levels (approx. 0.071 chance, velocities at ~0.161) when using velocity-only or mean-centered positional features. This demonstrates that decodable speaker identity is heavily dominated by static positional offsets that a 1D A-P warp preserves rather than removes.

Coordinate ConditionDispersion (mm)DDK LOSO F1/F2 (rr)Speaker ID Accuracy
Baseline31.280.387 / 0.6400.751
Scale-only26.360.381 / 0.6460.800
Scale+warp26.130.282 / 0.6120.902

Limitations

The study is limited by a small sample size (N=14N=14 speakers with complete multimodal data) drawn from a single language (German). The non-rigid mapping is restricted to a 1D anterior-posterior warp, which cannot capture complex 2D variations in palate curvature, vertical tongue shaping, or pharyngeal geometry. Additionally, acoustic pharyngometry measurements are static and sensitive to posture and mouthpiece protocol variations.

Why read this

Speech researchers and engineers working on cross-speaker kinematic normalization or articulatory-to-acoustic mapping should read this to understand why achieving geometric trajectory alignment does not inherently yield speaker-independent acoustic prediction.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Cross-speaker articulatory analysis, phonetic comparison, and phonetic corpus standardization.

Institutions

Zurich Forensic Science Institute, University of Zurich

Funding / 經費: Swiss National Science Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1883