TL;DR — GETS is a novel electromyography-to-speech (ETS) framework that integrates a diffusion-based generative network with silent speech recognition (SSR) semantic guidance, achieving a new state-of-the-art word error rate (WER) of 11.89% on silent EMG test data.
Key contributions
- Introduces GETS, setting a new SOTA WER of 11.89% on silent EMG test sets by fusing acoustic EMG cues with SSR-based semantic guidance.
- First application of diffusion-based classifier guidance combined with silent speech recognition to resolve underspecified EMG inputs without retraining the generative backbone.
- Provides extensive temporal alignment and energy contour analyses proving that the framework preserves underlying prosodic traits rather than acting as a naive text-to-speech system.
Problem
Silent speech interfaces aim to recover audible speech from facial electromyography (EMG) to assist individuals with speech or motor impairments, but direct electromyography-to-speech (ETS) models struggle with semantic intelligibility and yield high word error rates exceeding 25%. Prior paradigms like direct acoustic mapping, soft unit prediction (SU-ETS), and initial diffusion models (DiffETS) prioritize acoustic reconstruction over linguistic correctness, failing to disambiguate inherently underspecified silent EMG signals. This gap matters because minimizing simple acoustic loss does not guarantee semantic or linguistic consistency, creating an intelligibility bottleneck that limits real-world usability.
Method
The GETS framework consists of an EMG-conditioned mel generator and an SSR-guided semantic refinement mechanism. The core generator uses a DiffWave backbone combined with a shared EMG encoder containing convolutional and transformer layers. During training, Dynamic Time Warping (DTW) aligns asynchronous silent EMG features () to paired voiced EMG features (), generating temporally aligned features () that condition the DDPM forward and reverse processes alongside a classifier-free guidance weight .
At inference time, silent speech recognition (using pre-trained MONA-LISA ensembled with fine-tuned GPT-4) produces a predicted text transcript . A frozen audio-only ASR model (AVEC variant pre-trained on LRS2/LRS3) acts as a classifier evaluating conditional log-likelihoods log during reverse diffusion steps (). The gradients of this classifier are integrated into the noise estimate via classifier guidance scaled by weight with a normalization factor .
The model is trained using the Adam optimizer with an initial learning rate of , a batch size of 8 per A6000 GPU, and an loss over 400 epochs with diffusion steps and conditional dropout . Generated mel-spectrograms are converted to waveforms using a pre-trained HiFi-GAN vocoder.
Experimental setup
Evaluated on the Gaddy & Klein dataset consisting of 8-channel monopolar EMG signals sampled at 1 kHz (resampled to 800 Hz) and audio at 16 kHz, containing 1,285 silent-voiced parallel pairs and 5,470 non-parallel voiced training utterances from a single subject. Compared against baselines including Gaddy and Klein, SU-ETS, diff-ETS, transduction models, and optimized EMG encoder configurations. Metrics include Word Error Rate (WER evaluated via Whisper-medium and DeepSpeech), Mean Absolute Error (MAE) in milliseconds for word/pause onset and offset timing, and Root Mean Square Error (RMSE) for frame-level log-mel energy contours. Implemented on NVIDIA A6000 GPUs.
Results
GETS achieves a headline WER of 11.89% using Whisper-medium evaluation, slashing error rates by more than half compared to prior single-subject ETS baselines (e.g., Gaddy and Klein at 25.74%, SU-ETS at 26.29%, and diff-ETS at 32.1%). When evaluated via DeepSpeech, GETS attains a WER of 21.32%, outperforming Transduction (33.51%) and Optimized EMG Encoder models (35.4%). Ablation studies demonstrate that neutralizing EMG guidance () causes WER to spike to 18.81% (and up to 32.78% when both components degrade), while removing semantic guidance () degrades WER to 12.93%. Prosodic evaluations confirm strong temporal alignment and an energy contour RMSE of 0.437, which collapses heavily under EMG-shuffling (RMSE 1.185) compared to text-shuffling (RMSE 0.608), confirming that prosody is successfully anchored to the raw biosignals.
| Model | WER (%) (Whisper) | WER (%) (DeepSpeech) |
|---|---|---|
| Gaddy and Klein | 25.74 | - |
| SU-ETS | 26.29 | - |
| diff-ETS | 32.10 | - |
| Transduction Model | - | 33.51 |
| Optimized EMG Encoder | - | 35.40 |
| GETS (Ours) | 11.89 | 21.32 |
Limitations
The framework is evaluated exclusively on a single-subject dataset (Gaddy & Klein dataset), leaving multi-speaker generalization untested. The reliance on pre-trained silent speech recognition (MONA-LISA) and external LLM/ASR pipelines introduces cascading dependency risks, and the evaluation is limited to clean, controlled laboratory silent speech without real-world noise variations.
Why read this
Speech and ML researchers focusing on biosignal processing and guided diffusion should read this paper to learn how to inject classifier-based semantic guidance into non-acoustic generative pipelines without destroying underlying signal-derived prosody.
Code
Applications
Assistive communication devices for individuals with severe speech or motor impairments, and silent speech interfaces for secure or private communications.
Institutions
Seoul National University
Funding / 經費: National Research Foundation of Korea, Institute of Information & Communications Technology Planning & Evaluation, Ministry of Science and ICT, National IT Industry Promotion Agency
Related
- TAP-ETS: Time Aligned Phoneme Guiding for EMG-to-Speech Synthesis — same problem · relatedness 3.0/3
- Enhancing EMG-to-Speech via Silent-Voiced Representation Alignment — same problem · relatedness 2.9/3
- Tongue2Speech: Real-Time Speech Synthesis from Tongue Ultrasound Videos via Spatiotemporal Transformers — same problem · relatedness 2.1/3
- Speaker-Independent Speech Synthesis from Real-time MRI Articulatory Data — same problem · relatedness 2.1/3
- Emergence of Phonetic Representations in EMG-based Silent Speech Interfaces — same problem · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1938