All papers
Speech synthesisFull-paper digest

Tongue2Speech: Real-Time Speech Synthesis from Tongue Ultrasound Videos via Spatiotemporal Transformers

Yash Sonkar, Yasaswi Kilaru, Sudheera Yelimeli, Neil Shah, Vineet Gandhi

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.6 KB · Ready to paste

Preview copied content

TL;DR — Tongue2Speech is a lightweight non-autoregressive framework that maps tongue ultrasound video directly to mel-spectrograms using 3D spatio-temporal convolutions and a Transformer encoder. It achieves a state-of-the-art 15.93% WER in single-speaker settings and 31.64% WER in multi-speaker settings on the TaL corpus.

Key contributions

  • Proposes a lightweight non-autoregressive U2S architecture with 19.17M total parameters (6.25M excluding vocoder) combining a 3D CNN front-end and a Transformer encoder.
  • Demonstrates that 3D spatial-temporal encoding with self-attention drastically outperforms legacy recurrent or purely convolutional architectures for modeling articulatory coarticulation.
  • Comprehensively evaluates on the TaL corpus across single-speaker and multi-speaker settings, proving that ASR-based WER correlates with intelligibility better than spectral MSE.
  • Shows that zero-shot multi-speaker generalization remains difficult, but fine-tuning on limited speaker-specific data successfully restores intelligibility for unseen users.

Problem

Silent speech interfaces (SSIs) are vital for individuals with lost phonation caused by laryngectomy or neurological disorders, but prior methods rely on invasive intraoral sensors (EMA), expensive high-cost real-time MRI, or indirect muscle recordings (sEMG). Ultrasound tongue imaging (UTI) directly captures dynamic tongue surface geometry with portable hardware, yet prior UTI-to-speech models rely on convolutional-recurrent (LSTM) backbones that struggle with long-range coarticulation. Furthermore, prior literature depends heavily on weak spectral distortion metrics like MSE and MCD which do not track lexical intelligibility, and lacks large-scale multi-speaker and word error rate evaluations.

Method

Tongue2Speech takes raw ultrasound scanline frames in polar coordinates (64 scanlines × 842 radial samples) and converts them to a Cartesian wedge representation of size 64 × 64 pixels per frame using GPU-accelerated bilinear grid sampling. The wedge sequences pass through four stacked 3D convolutional blocks (Conv3D + BatchNorm + ReLU) that jointly operate across spatial and short-term temporal dimensions, followed by adaptive average pooling along spatial dimensions to yield per-frame 256-dimensional latents.

This latent sequence is fed into a 6-layer encoder-only Transformer with 8 attention heads per layer, a 1024 feed-forward hidden dimension, GELU activations, and sinusoidal positional encodings. The resulting hidden states are projected via a 2-layer MLP to 80-dimensional mel-spectrograms, trained using L1 loss. A HiFi-GAN neural vocoder then synthesizes the final waveform audio from the predicted mel-spectrograms.

Experimental setup

Evaluated on the TaL corpus (TaL1 single-speaker across 5 sessions, and TaL80 multi-speaker across 81 speakers yielding ~24 hours of synchronized data). Compared against 6 baselines: Conformer-U2S, STN-CNN, 3D-CNN, 3D-CNN+BiLSTM, 3D-CNN+BiLSTM w/ Skip, and 2D-CNN+BiLSTM. Evaluated via OpenAI Whisper Medium ASR Word Error Rate (WER) and log-mel MSE. Multi-speaker personalization uses 1500 epochs of pre-training on 77 speakers, followed by 20 epochs of fine-tuning on 95% of data for held-out speakers.

Results

Tongue2Speech achieves an overall single-speaker WER of 15.93% and MSE of 0.594 on TaL1, outperforming the strongest baseline (3D-CNN+BiLSTM w/ Skip at 24.15% WER, 0.793 MSE) by over 8 absolute WER points. Re-implemented Conformer-U2S and STN-CNN collapse on this task with near-saturated WERs near ~99% and high MSEs (1.8-1.9). In the multi-speaker TaL80 evaluation, Tongue2Speech achieves 31.64% WER versus 42.46% for standard 3D-CNN+BiLSTM and 33.23% for the skip variant. Speaker-specific fine-tuning on unseen held-out speakers successfully drops WER down to 27.62% for Speaker 80, demonstrating effective personalization.

SystemTaL1 WER (%)TaL1 MSETaL80 WER (%)TaL80 MSE
Conformer-U2S99.261.819--
STN-CNN99.081.877--
2D-CNN + BiLSTM60.291.417--
3D-CNN + BiLSTM36.460.90642.460.981
3D-CNN + BiLSTM w/ Skip24.150.79333.230.863
Tongue2Speech (Ours)15.930.59431.640.851

Limitations

Zero-shot cross-speaker generalization remains difficult due to anatomical variability and probe placement differences, requiring speaker-specific fine-tuning. Evaluation is constrained to English speakers from the TaL corpus, leaving multilingual scalability unverified. Session-level variability (up to 9% WER fluctuations across days) demonstrates sensitivity to probe placement shifts.

Why read this

Researchers and engineers building silent speech interfaces or articulatory-to-acoustic inversion models should read this paper to see how replacing recurrent networks with a lightweight 3D-CNN and Transformer backbone substantially boosts lexical intelligibility. It also provides a cautionary blueprint proving that spectral distortion metrics like MSE fail to indicate actual speech intelligibility.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Practical silent speech interfaces for individuals with voice impairments (e.g., laryngectomy patients) and auxiliary articulatory feature conditioning for speech recognition.

Institutions

International Institute of Information Technology Hyderabad, TCS Research

Funding / 經費: Anusandhan National Research Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3023