All papers
Speech recognitionFull-paper digest

SEA-MDD: Self-adapting Mispronunciation Detection and Diagnosis Models via Test-Time Training

Minglin Wu, Helen Meng

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.2 KB · Ready to paste

Preview copied content

TL;DR — SEA-MDD introduces a self-adapting mispronunciation detection and diagnosis framework via Test-Time Training (TTT), dynamically updating model weights on a single test sentence to achieve an 8.03% phoneme error rate and an 81.54% F1 score on L2 English speech.

Key contributions

  • Proposes SEA-MDD, integrating multilayer perceptron (MLP)-based Test-Time Training (TTT) modules into wav2vec 2.0 Transformer blocks for single-utterance speech adaptation.
  • Eliminates the heavy data requirements of prior meta-learning adaptation methods, adapting successfully using only a single incoming test sentence.
  • Achieves consistent performance gains over wav2vec2-CTC and MAML baselines, yielding a lower phoneme error rate (8.03%) and higher F1 score (81.54%).
  • Provides extensive empirical ablations over TTT module positions (early vs. late layers) and architectures (Linear vs. MLP).

Problem

Mispronunciation detection and diagnosis (MDD) systems for second language (L2) learners suffer from severe performance degradation when exposed to speakers and error types that diverge from fixed training distributions. Collecting and annotating large-scale L2 corpora covering all possible mispronunciation variations is prohibitively expensive. Prior adaptation approaches like model-agnostic meta-learning (MAML) require substantial target-speaker adaptation data (e.g., hours of speech), leading to a secondary data scarcity bottleneck.

Method

The SEA-MDD framework builds on the wav2vec 2.0 base model, which features 12 Transformer blocks (768 model dimension, 3072 feed-forward dimension, 8 attention heads) processing latent features from a 7-layer temporal convolutional encoder. A TTT module is inserted into the Transformer blocks prior to the feed-forward sub-layer. The TTT module is implemented as either a single linear layer (TTT-Linear) or a two-layer MLP with a GELU activation in between (TTT-MLP), followed by layer normalization, a residual connection, and a learnable gating mechanism (Gate(x)=tanh⁡(α)⊙xGate(x) = \tanh(\alpha) \odot x).

During both training and test time, the parameters WW of the TTT module undergo an inner-loop update via a single gradient descent step on a self-supervised reconstruction task. The input speech representation xtx_t is projected via θK\theta_K into keys and via θV\theta_V into reconstruction targets, while a query projection θQ\theta_Q is applied to compute the output. The self-supervised loss L\mathcal{L} minimizes the squared error between the MLP output and the reconstruction target θVxt\theta_V x_t, with a learning rate η=1e−3\eta = 1\text{e}-3. A mini-batch strategy (b=32b=32) is used along the temporal dimension for efficiency. The outer loop optimizes initial parameters W0W_0 and all other network weights using CTC loss, the tri-stage learning rate schedule (peak lr 5e−45\text{e}-4), the Adam optimizer, for 20,000 updates on two NVIDIA H100 GPUs.

At inference time, the model executes only the inner-loop weight update using the current test sentence, allowing zero-shot dynamic adaptation without auxiliary target-speaker datasets. Triton kernels, data loading-computation asynchrony, and sequence-dimension gradient checkpointing are utilized to maintain low adaptation latency.

Experimental setup

Evaluated on the CU-CHLOE L2 English dataset (34.6 hours total across 210 Cantonese and Mandarin speakers, split into 24 hours training, 3.6 hours validation, and 7 hours test). Compared against non-adaptable wav2vec2-CTC and speaker-adaptable wav2vec2-MAML (which requires 2 hours of target-speaker adaptation data). Metrics include Phoneme Error Rate (PER), False Rejection Rate (FRR), False Acceptance Rate (FAR), Precision, Recall, F1 score, and Diagnosis Accuracy (DIAA).

Results

SEA-MDD-MLP (All blocks) achieves the headline performance with a PER of 8.03%, FRR of 4.40%, FAR of 19.62%, Precision of 82.72%, Recall of 80.38%, F1 score of 81.54%, and DIAA of 94.06%, outperforming the wav2vec2-CTC baseline (PER 8.53%, F1 80.40%) and wav2vec2-MAML (PER 8.46%, F1 80.67%). Inserting TTT modules into all 12 blocks outperforms single-block insertion (e.g., SEA-MDD-MLP 1st block yields 8.18% PER and 81.18% F1), and MLP variants consistently outperform Linear variants.

In computational overhead comparisons, wav2vec2-MAML demands 94.4M parameters, 30 minutes of latency, and 2 hours of data, whereas SEA-MDD-MLP (1st block) uses only 3.0M parameters, 5 ms latency, and a single sentence. However, placing TTT modules deeper than layer index 5 shows diminishing returns, indicating that early-layer feature adaptation is crucial.

SystemsPER (%) ↓FRR (%) ↓FAR (%) ↓Precision (%) ↑Recall (%) ↑F1 (%) ↑DIAA (%) ↑
wav2vec2-CTC8.534.5920.9781.8279.0380.4093.71
wav2vec2-MAML8.464.5720.6382.0179.3780.6793.63
SEA-MDD-Linear (1st block)8.264.4320.5082.6279.5081.0393.83
SEA-MDD-Linear (All blocks)8.134.4120.0282.6279.9881.2893.87
SEA-MDD-MLP (1st block)8.184.4520.0882.4879.9281.1893.89
SEA-MDD-MLP (All blocks)8.034.4019.6282.7280.3881.5494.06

Limitations

The evaluation is restricted to L2 English speech from Cantonese and Mandarin native speakers within a single dataset (CU-CHLOE). The approach has not yet been validated across broader L1 accent backgrounds, varying noise or recording conditions, or larger foundation speech models beyond wav2vec 2.0 base.

Why read this

Speech researchers and engineers working on computer-assisted pronunciation training will learn how test-time training can eliminate multi-speaker fine-tuning data bottlenecks for mispronunciation detection and diagnosis.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Computer-assisted language learning (CALL) software, automated pronunciation scoring and tutoring systems, and real-time L2 speech feedback tools.

Institutions

Chinese University of Hong Kong

Funding / 經費: Centre for Perceptual and Interactive Intelligence, Innovation and Technology Commission of the Hong Kong Special Administrative Region Government

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-856