All papers
Health & clinical speechFull-paper digest

Formant-Guided Speech Repair for Enhanced Comprehension of Dysarthric Speech

Xin-Yu Chen, Jing-Tong Tzeng, Carlos Busso, Chi-Chun Lee

Code & resourcesgithub.com/xinyu0308/FAST-SR

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.9 KB · Ready to paste

Preview copied content

TL;DR — The paper introduces Formant-Aligned Speech Repair (FAST), a modular speech-in, speech-out framework that explicitly regularizes distorted vowel formants prior to neural TTS synthesis, achieving a 72.4% relative reduction in character error rate on Mandarin dysarthric speech.

Key contributions

  • Proposes the Formant-Aligned Spectral Transformation (FAST) module that explicitly remaps distorted vowel formants toward healthy reference distributions using a Gaussian-weighted warping function.
  • Combines a dysarthria-adapted ASR (Wav2Vec 2.0 + Conformer) with speaker-adaptive TTS (XTTS v2 backend with a fine-tuned speaker encoder) for speech reconstruction.
  • Demonstrates robust cross-lingual and cross-pathology generalization across four Mandarin and English dysarthric speech datasets (CDSD, MDSC, MSDM, TORGO).
  • Achieves lower character error rates than Oracle TTS on certain corpora, proving that explicit acoustic vowel space regularization is as critical as correct textual conditioning.

Problem

Dysarthria severely degrades speech intelligibility due to impaired articulatory control that compresses the vowel space and eliminates spectral contrast. Existing automatic speech recognition systems only output text, thereby discarding vital paralinguistic and prosodic cues essential for natural human interaction. Meanwhile, prior speech-in, speech-out neural reconstruction methods like voice conversion or latent diffusion models operate as black boxes that struggle with severe non-linear articulatory distortions and fail to achieve high intelligibility.

Method

The system processes source dysarthric speech through three cascaded modules: a dysarthria-adapted ASR, a formant correction module, and a speaker-adaptive TTS. The ASR module uses frozen Wav2Vec 2.0 front-ends (WenetSpeech for Mandarin, LibriSpeech for English) feeding a Conformer encoder and recurrent decoder, trained on a hybrid CTC and attention cross-entropy objective with lambda_CTC set to 0.3.

The FAST module obtains phone boundaries via the Montreal Forced Aligner, estimates F1 and F2 across the central 80% of vowel segments using Praat's Burg method, and computes deviations against healthy reference values derived from AISHELL-1 (Mandarin) and TORGO (English) stratified by vowel class (/a/, /e/, /i/, /o/, /u/) and gender. It applies a Gaussian-weighted spectral warping function governed by a repair factor scaling parameter kappa (set optimally to 2.0), followed by inverse STFT, spectral cross-fading, RMS energy normalization, and high-frequency enhancement.

The Speaker-Adaptive TTS is built on the XTTS v2 framework, where the core backbone is frozen and only the speaker encoder is fine-tuned to map utterance embeddings close to target speaker centroids. The TTS is conditioned on the FAST-corrected speech waveform for prosody and timbre while taking ASR-predicted text transcripts as linguistic input.

Experimental setup

Evaluated on four pathological speech datasets: CDSD (Mandarin, 29.4h, 44 speakers), MDSC (Mandarin, 9.1h, 21 speakers), MSDM (Mandarin, 2.2h, 62 speakers), and TORGO (English, 2.5h, 8 speakers), covering cerebral palsy, stroke, ALS, and degeneration. Compared against baseline systems including DiffDSR, RnV, and Liu et al., plus ablation variants (w/o SPK, w/o FAST). Metrics include Character Error Rate (CER), Word Error Rate (WER), UTMOS naturalness, speaker cosine similarity via Resemblyzer, Formant Centralization Ratio (FCR), Vowel Space Area (VSA), and a 5-point subjective MOS evaluation with 18 native listeners.

Results

The full proposed system achieves a CER of 22.33% on CDSD and 38.44% on MDSC, representing massive error reductions compared to input speech (80.97% and 88.92%) and outperforming external baselines like DiffDSR (95.78% CER on CDSD) and RnV. On the TORGO English corpus, it achieves 13.65% WER compared to 32.06% for the raw input. Ablations show that removing FAST or speaker adaptation degrades CER/WER by 4-7% absolute. Subjective evaluations demonstrate over 90% gains in intelligibility, comprehension, and fluency, alongside a 60.9% reduction in listening effort (dropping from 3.81 to 1.49). Speaker similarity exhibits a slight trade-off, dropping from around 0.74-0.76 down to 0.63-0.68 in exchange for large intelligibility improvements.

System / ConditionCDSD (CER %)MDSC (CER %)MSDM (CER %)TORGO (WER %)UTMOS (CDSD)
Input (Dysarthric)80.9788.9248.6532.061.576
Input + FAST69.9176.9544.9832.901.546
ASR + TTS29.9546.9136.8420.672.083
Ours (w/o FAST)28.5245.3835.6023.182.076
Ours (Full)22.3338.4434.5013.652.250
Oracle TTS23.8039.7432.3889.652.061

Limitations

The framework relies on accurate phone alignments and text transcripts from ASR to perform formant correction, meaning severe ASR failures can propagate errors to the FAST module. Articulatory repair is currently restricted to vowels, omitting consonants which also contribute significantly to dysarthric speech degradation. Furthermore, the system experiences a minor trade-off wherein enhanced intelligibility slightly reduces native speaker voice similarity.

Why read this

Researchers and engineers building speech-to-speech assistive communication systems will learn how to integrate linguistically grounded acoustic rules (formant correction) with neural TTS generative priors to bypass the limitations of black-box diffusion or voice conversion models.

Code

Applications

Assistive communication devices for individuals with motor speech disorders, smart-home voice assistants adapted for pathological speech, and real-time conversational speech repair interfaces.

Institutions

National Tsing Hua University, Carnegie Mellon University

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1217