All papers
Health & clinical speechFull-paper digest

Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech

Jaesung Bae, Xiuwen Zheng, Minje Kim, Chang D. Yoo, Mark Hasegawa-Johnson

Code & resourcesgithub.com/JaesungBae/DA-DSQA

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.8 KB · Ready to paste

Preview copied content

TL;DR — A three-stage framework leveraging pseudo-labeling, contrastive representation learning, and external typical speech achieves robust dysarthric speech quality assessment (DSQA), reaching an average SRCC of 0.761 on unseen cross-domain test sets. By combining 232.9 hours of unlabeled dysarthric data (Speech Accessibility Project) with 921.7 hours of clean typical speech (LibriSpeech) under a coarse binary contrastive objective, it substantially outperforms existing SQA baselines.

Key contributions

  • A three-stage framework (pseudo-labeling, weakly supervised contrastive pretraining, and fine-tuning) that fully exploits large quantities of unlabeled dysarthric speech.
  • Integration of a large-scale typical speech corpus (LibriSpeech) to dramatically expand acoustic and speaker environment diversity.
  • A coarse binary grouping contrastive loss strategy (L_coarse) that effectively bridges the distribution gap between typical and pathological speech without overfitting to noisy pseudo-labels.
  • Comprehensive cross-domain evaluation across five diverse datasets spanning multiple languages (English, Mandarin, Italian, Czech, Slovak, Spanish) and clinical etiologies.

Problem

Dysarthric speech quality assessment (DSQA) is crucial for clinical monitoring and inclusive technologies like ASR and speech enhancement, but collecting expert ratings from speech-language pathologists is expensive and difficult to scale. Existing models trained purely on small labeled subsets or non-pathological speech struggle to generalize across diverse clinical etiologies, acoustic environments, and languages. Prior speech foundation models and non-intrusive metrics (e.g., DNSMOS, UTMOS) fail to reliably capture the perceptual and intelligibility dimensions unique to pathological speech.

Method

The system relies on a frozen Whisper-large-v3 encoder to extract frame-level features, followed by two linear projection layers, statistical temporal average pooling, and a final prediction layer. In Stage 1, a regression model is trained on 10.8 hours of labeled Speech Accessibility Project (SAP) data using Huber loss (delta = 0.5) to generate pseudo-labels for 232.9 hours of unlabeled SAP speech. In Stage 2, weakly supervised contrastive pretraining is performed using the pseudo-labeled SAP data, labeled SAP data, and 921.7 hours of LibriSpeech (assigned a typical label of 1). Three contrastive strategies are evaluated (discrete, continuous distance thresholding, and coarse binary thresholding L_coarse where beta = 1.5), alongside VICReg variance regularization (gamma = 1.0) to prevent feature collapse. Larger temperatures (tau = 10.0 for L_coarse) prevent dataset separation between LibriSpeech and SAP.

In Stage 3, the pretrained adaptation layers are frozen/initialized, a fresh linear regression head is added, and the model is fine-tuned end-to-end on the labeled SAP dataset using AdamW with a learning rate of 1e-4 and label-weighted random sampling for 10 epochs. The entire architecture decouples robust representation learning from final fine-grained regression, mitigating label noise while adapting representations to clinical severity.

Experimental setup

Evaluated on the English Speech Accessibility Project (SAP: 10.8 hours labeled, 232.9 hours unlabeled for training; 3.1 hours for testing) and LibriSpeech (921.7 hours). Cross-domain zero-shot evaluation is performed on five multilingual datasets: UASpeech (English, 7.8 hrs), DysArinVox (Mandarin, 2.3 hrs), EasyCall (Italian, 10.0 hrs), EWA-DB (Czech/Slovak, 4.4 hrs), and NeuroVoz (Spanish, 1.7 hrs). Baselines include DNSMOS, UTMOS, SpICE, HuBERT Probe, standard fine-tuned Whisper (Baseline), SimCLR, and Rank-N-Contrast (RNC). Metrics are Spearman's Rank Correlation Coefficient (SRCC) and Pearson Correlation Coefficient (PCC).

Results

The proposed L_coarse framework achieves an average cross-domain SRCC of 0.761 (PCC 0.749) and an in-domain SAP SRCC of 0.716, significantly outperforming DNSMOS (0.105 avg SRCC), UTMOS (0.346 avg SRCC), SpICE (0.450 avg SRCC), and HuBERT Probe (0.621 avg SRCC). Compared to the raw Whisper Baseline (0.732 cross-domain SRCC), L_coarse improves cross-domain robustness across nearly all test sets, hitting 0.975 SRCC on UASpeech and 0.631 on DysArinVox. Ablations show that dropping LibriSpeech data reduces cross-domain average SRCC to 0.734, while omitting variance regularization or using fine-grained pseudo-labeling hurts out-of-domain generalization due to overfitting.

SystemSAP (In-Domain) SRCCUASpeech SRCCDysArinVox SRCCEasyCall SRCCEWA-DB SRCCNeuroVoz SRCCCross-Domain Avg SRCC
DNSMOS0.1860.7500.3700.041-0.274-0.3610.105
UTMOS0.4890.9620.521-0.0510.328-0.0280.346
SpICE0.4730.9360.5380.2050.3930.1800.450
HuBERT Probe0.5310.9270.2790.6040.5880.7050.621
Baseline0.7190.9490.5780.8490.7090.5750.732
Proposed (L_coarse)0.7160.9750.6310.8720.7110.6170.761

Limitations

Training relies entirely on English-only dysarthric and typical datasets, limiting direct supervision scope despite zero-shot cross-lingual evaluations. The approach depends on initial pseudo-labels generated by a model trained on a small labeled fraction (3% of SAP), which can constrain the ceiling of representations if initial model bias is severe. Evaluation relies solely on utterance-level aggregations compared against speaker-level clinical scales, leaving fine-grained temporal error localization unexplored.

Why read this

Read this paper to learn how to effectively combine limited domain-specific labeled data, large unlabeled target data, and auxiliary out-of-domain typical corpora via weakly supervised contrastive learning. It offers a masterclass in adapting robust speech foundation models (Whisper) for pathological and low-resource acoustic regression tasks.

Code

Applications

Automated clinical screening of motor speech disorders, continuous speech rehabilitation monitoring, and data filtering/augmentation for pathological automatic speech recognition (ASR) and speech synthesis.

Institutions

University of Illinois Urbana-Champaign, Korea Advanced Institute of Science & Technology

Funding / 經費: Institute of Information & Communications Technology Planning & Evaluation, National Science Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1390