All papers
Phonetics & linguisticsFull-paper digest

Towards Language-Agnostic Speech Inversion

Saba Tabatabaee, Mark Tiede, Suzanne Boyce, Liran Oren, Carol Espy-Wilson

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.0 KB · Ready to paste

Preview copied content

TL;DR — This paper presents a multi-task speech inversion system that maps raw audio to vocal tract variables, source features, and velopharyngeal motion using a pretrained WavLM-Large backbone, achieving strong cross-lingual generalizability on untrained French and Russian speakers despite being trained exclusively on English data.

Key contributions

  • Collected a multi-lingual dataset comprising co-recorded electromagnetic articulography (EMA), nasalance, and speech audio data.
  • Performed cross-lingual evaluations of a speech inversion system estimating oral tract variables and source features on unseen languages.
  • Extended the speech inversion evaluation framework to predict velopharyngeal tract variables (nasalance proxies) across languages.
  • Compared proposed speech inversion multi-task performance against prior baselines on standard English articulatory benchmarks.

Problem

Recovering articulatory timing and spatial patterns directly from speech acoustics (speech inversion) is challenging due to acoustic interaction effects and many-to-one mapping between vocal tract shapes and acoustic outputs. Prior speech inversion systems are almost exclusively developed and evaluated on English-language datasets, leaving their cross-linguistic generalizability largely unverified. Furthermore, collecting direct articulatory data requires expensive and specialized equipment like electromagnetic articulography or X-ray microbeam, making scalable data collection difficult, especially for under-resourced languages and vulnerable populations.

Method

The architecture takes input speech signals and extracts representations from all 25 hidden layers of a pretrained WavLM-Large model, computing a weighted sum of layer-wise embeddings. This unified representation is passed through 3 Conformer layers to capture local and long-range temporal dependencies. The Conformer output feeds into a fully connected layer with 256 hidden units and a GELU activation, followed by a second fully connected layer with 128 hidden units. Because WavLM embeddings operate at 50 Hz while target articulatory outputs are sampled at 100 Hz, an upsampling block by a factor of two with batch normalization is applied.

Multi-task learning is implemented via separate output dense layers: one predicting six oral tract variables (lip aperture, lip protrusion, tongue body constriction location/degree, tongue tip constriction location/degree) and another predicting three source features (periodicity, aperiodicity, and fundamental frequency) or jointly predicting source features plus a velopharyngeal tract variable. Training uses the AdamW optimizer with an initial learning rate of 5e-4, weight decay of 1e-3, batch size of 8, a plateau-based learning rate scheduler (patience of 5 epochs), and early stopping (patience of 8 epochs). The loss function combines Pearson correlation and root mean square error with alpha set to 0.2.

The system is trained on a combination of the XRMB English dataset (36 speakers, 5.74 hours) and the YU English dataset (12 speakers, 3.21 hours), and evaluated zero-shot on French (4 speakers, 59 minutes) and Russian (3 speakers, 35 minutes) subsets.

Experimental setup

Evaluated on the XRMB dataset (5.74 hours across 46 English speakers) and the YU dataset (7.17 hours across 27 speakers spanning English, French, and Russian). Baselines include the prior English speech inversion model from Tabatabaee et al. (2024/2026). Performance is measured using Pearson product-moment correlation (PPMC) scores between estimated trajectories and ground-truth sensor or nasalance measurements.

Results

On the XRMB test set, the proposed model achieves an average PPMC score of 0.86 across all nine parameters (vs 0.85 for the baseline). On the YU English test set, it achieves an average PPMC of 0.85. When evaluated zero-shot on unseen languages, the model achieves average PPMC scores of 0.83 for French and 0.74 for Russian. For velopharyngeal (VP) tract variable estimation against ground-truth nasalance, the model achieves PPMC scores of 0.92 on English, 0.89 on French, and 0.82 on Russian. Performance drops slightly on Russian pitch/periodicity due to a noisier recording environment for one speaker.

SystemLanguageLALPTBCLTBCDTTCLTTCDPerAperF0Avg All
SI model [8]English (XRMB)0.910.760.800.860.840.950.940.880.750.85
Proposed SIEnglish (XRMB)0.930.760.780.870.820.950.950.900.790.86
Proposed SIEnglish (YU)0.870.860.810.840.820.870.950.830.840.85
Proposed SIFrench (YU)0.880.870.800.830.780.840.950.800.760.83
Proposed SIRussian (YU)0.760.750.700.710.690.710.910.680.710.74

Limitations

Evaluated on a limited number of speakers for non-English languages (4 French speakers, 3 Russian speakers, and only 1 Russian speaker for VP evaluation). The training data scale is relatively small (under 13 total hours across XRMB and YU). Acoustic variations such as background noise heavily degrade pitch and periodicity estimation accuracy on specific speakers.

Why read this

Speech and ML researchers building cross-lingual or articulatory speech representations should read this to see how self-supervised speech models like WavLM capture universal human vocal tract dynamics transferable across languages without fine-tuning.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Multi-lingual speech processing, computer-assisted pronunciation training, speech therapy, and clinical assessments of speech and swallowing disorders.

Institutions

University of Maryland, Yale University, University of Cincinnati

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1633