TL;DR — This paper proposes a confidence-score guided incremental and speaker-adaptive pseudo-labeling framework for semi-supervised elderly speech recognition, achieving statistically significant WER/CER reductions of 1.45% and 2.27% absolute on English and Cantonese elderly datasets.
Key contributions
- A curriculum learning strategy that partitions unlabeled data from high to low confidence per speaker, avoiding error accumulation and complete utterance discard.
- A lightweight Confidence Estimation Module (CEM) built from a 3-layer residual feedforward network with a sigmoid gate to produce token- and utterance-level reliability ranks.
- Integration of speaker prompt-adapted LoRA-Whisper models into the iterative updates to mitigate severe speaker heterogeneity in elderly and clinical populations.
- Demonstration of robust performance matching fully supervised baselines using only 10% labeled data across DementiaBank Pitt and JCCOCC MoCA corpora.
Problem
Elderly and clinical speech exhibits high heterogeneity, such as imprecise articulation from neuromotor decay and linguistic degradation from cognitive decline, causing high error rates in standard speech foundation models. Existing semi-supervised methods suffer from three key flaws: confidence-based filtering entirely discards low-confidence utterances (hurting speaker modeling), prior incremental methods lack reliability-based partitioning which leads to error cascades, and traditional setups rely on speaker-independent models that fail to capture individual physiological variations. This makes effective exploitation of large pools of untranscribed clinical tele-consultation data extremely difficult.
Method
The method uses Whisper-medium as the base architecture, processing log-Mel spectrogram inputs downsampled via a convolutional block, with LoRA parameters (rank 8, alpha 16) applied to attention query, key, value, and output layers. A lightweight Confidence Estimation Module (CEM)—a 3-layer residual feedforward network with a hidden dimension of 64, batch normalization, dropout, and a sigmoid gate—takes decoder token outputs concatenated with top-10 score logits to predict confidence against edit-distance-aligned targets (1 for correct, 0 for substitution/insertion).
Unlabeled data is ranked per-speaker by utterance-level confidence (mean token confidence excluding special tokens) and uniformly partitioned into subsets from high to low confidence. Training follows a curriculum where model (fine-tuned on 10% labeled data ) decodes the highest-confidence subset . For subsequent iterations , model regenerates pseudo-labels for subset , which is merged with the labeled data and all previously accumulated pseudo-labeled subsets to update the model.
To capture speaker heterogeneity, speaker adaptive training (SAT) is incorporated by concatenating a set of learnable speaker prompts to the convolutional speech hidden states before Transformer encoding. During test-time adaptation (TTA) or speaker adaptive relabeling (SA relab.), speaker prompts are optimized for test speakers using generated pseudo-labels to iteratively refine downstream text transcriptions.
Experimental setup
Experiments use the DementiaBank Pitt corpus (33 hours, 292 AD assessment interviews, expanded to 58.9 hours via silence stripping/augmentation; 119 dev speakers, 95 eval speakers) and the Cantonese JCCOCC MoCA corpus (32.4 hours, 256 interviews, expanded to 156.9 hours; 49 dev and eval speakers). 10% of training speakers are randomly selected as labeled data with zero speaker overlap between labeled and unlabeled sets. Comparisons include fully supervised models, standard semi-supervised baselines, random sampling, and random partition incremental learning. Models are evaluated using Word Error Rate (WER%) and Character Error Rate (CER%) across participant (Par.) and clinical investigator (Inv.) subsets.
Results
The proposed incremental speaker-adaptive pseudo-labeling method outperforms the standard semi-supervised baseline (Sys. 3) by absolute WER/CER reductions of 1.45% and 2.27% (6.21% and 6.98% relative) on the DementiaBank Pitt and JCCOCC MoCA test sets. Ablations demonstrate that retaining 80% of pseudo-labeled data across 4 to 5 incremental steps yields optimal performance, while replacing confidence ranking with random partitioning degrades results. Under extreme low-resource conditions using only 5% labeled speakers, the method achieves a competitive WER of 23.43% on DementiaBank Pitt.
| System | Dataset / Condition | WER (%) / CER (%) |
|---|---|---|
| 100% Fully Supervised | DementiaBank Pitt (Eval All) | 20.43 |
| 10% Fully Supervised Sup. Baseline | DementiaBank Pitt (Eval All) | 23.70 |
| Standard Semi-Sup. Baseline | DementiaBank Pitt (Eval All) | 23.62 |
| Incremental SI (Conf. Score) | DementiaBank Pitt (Eval All) | 22.23 |
| Incremental SAT + SA Relab. | DementiaBank Pitt (Eval All) | 21.50 |
| Incremental SAT + SA Relab. | JCCOCC MoCA (Eval All) | 30.25 |
Limitations
The approach assumes speaker identities are known for unlabeled data to initialize and accumulate speaker prompts, which may not hold in completely anonymous clinical streams. The study focuses exclusively on elderly pathological assessment interviews (DementiaBank and MoCA), meaning generalizability to other noisy acoustic domains or general elderly conversational audio requires further validation. Compute requirements scale linearly with the number of iterative pseudo-labeling steps and prompt adaptation passes.
Why read this
Speech researchers and engineers building automatic speech recognition systems for medical, clinical, or elderly populations will find this a definitive recipe for turning unannotated patient interviews into high-performing models via curriculum-based pseudo-labeling and speaker prompt adaptation.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Automated speech recognition for elderly care, dementia screening interview transcription, tele-health clinical documentation, and low-resource medical dictation.
Institutions
Chinese University of Hong Kong, Chinese Academy of Sciences, National Research Council Canada
Funding / 經費: Hong Kong RGC GRF, Basic Research Project of Institute of Software, Chinese Academy of Sciences, Youth Innovation Promotion Association CAS
Related
- WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations — same problem · relatedness 2.0/3
- Two-stage semi-supervised learning with pseudo-labels: A case study on Northern Sámi ASR — shared technique · relatedness 2.0/3
- Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition — same problem · relatedness 2.0/3
- Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling — same problem · relatedness 2.0/3
- Conflict-Aware Pseudo-Labeling via Acoustic Signals for Multi-Task Speech Emotion Recognition — shared technique · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1611