TL;DR — This paper investigates discrete diffusion language models—specifically Masked Diffusion Language Models (MDLM) and Uniform-State Diffusion Models (USDM)—for ASR rescoring and proposes a novel token-level joint CTC-USDM decoding framework, achieving a WER of 4.47% on LibriSpeech dev-other while reaching an RTF of 0.003.
Key contributions
- Systematic comparison of Masked Diffusion Language Models (MDLM) and Uniform-State Diffusion Models (USDM) for ASR hypothesis rescoring.
- Proposed sample-level and global mask normalization techniques for MDLM rescoring that outperform naive sequence-length normalization.
- Designed a novel token-level joint decoding method combining frame-wise CTC probabilities with label-wise USDM distributions at each denoising step.
- Achieved substantial WER improvements with a single joint-decoding forward step while preserving near-real-time efficiency (RTF 0.003).
Problem
Traditional autoregressive language models used for ASR rescoring or joint decoding are constrained by a strictly left-to-right decoding structure, creating speed bottlenecks during parallel generation. While non-autoregressive discrete diffusion models (such as MDLM and USDM) bypass this restriction through bidirectional context and parallel text generation, their use as standalone standalone language models for token-level ASR joint decoding and structured rescoring remains underexplored. Applying them effectively requires overcoming high-variance ELBO estimators in rescoring and aligning continuous token distributions with frame-level acoustic models.
Method
The authors evaluate two Diffusion Transformer (DiT) architectures: a small 12-layer model (hidden dim 768, ~110M params) and a primary 24-layer model (16 attention heads, dropout 0.1, hidden dim 1024, ~340M params). Text is tokenized using SentencePiece into 10,240 subword units. MDLM corrupts text via independent masking based on a noise schedule αt, trained using a cross-entropy loss weighted over masked tokens. To fix high-variance Monte Carlo score estimation, the authors introduce sample-level mask normalization (normalizing each sample by its own mask count) and coupled scoring (creating complementary mask pairs).
USDM corrupts tokens via uniform vocabulary sampling rather than a mask token, yielding a full vocabulary probability distribution at every position during denoising. This property allows the formulation of a joint CTC-USDM decoding framework. At each denoising step l, token distributions Pθ,i are combined with frame-level CTC probabilities P_CTC,τi (renormalized over non-blank vocabulary) aligned via the first frame corresponding to each collapsed token. Ancestral sampling is then used to draw inputs for the next denoising step from the resulting softmax-normalized combined distribution.
Experimental setup
Experiments use LibriSpeech: the CTC ASR model is trained on 960 hours of training data, while the Diffusion LMs are trained on combined LibriSpeech LM data and train-other transcriptions. Models are optimized using AdamW (weight decay 0.1), a piecewise linear learning-rate scheduler, and a batch size of 20,000 tokens for 5, 10, and 25 epochs. Evaluation is conducted on LibriSpeech dev-other using Word Error Rate (WER) and Real-Time Factor (RTF), comparing against greedy CTC, autoregressive LM rescoring/first-pass decoding, and unnormalized diffusion baselines.
Results
MDLM rescoring with sample-level mask normalization reduces the CTC baseline WER of 5.08% down to 4.47% at K = 256 (25 training epochs). USDM rescoring reaches 4.71% WER with 24 layers at K = 256. Meanwhile, the proposed CTC-USDM joint decoding achieves 4.66% WER using only a single denoising step (L = 1, t_start = 0.1) with an extremely fast RTF of 0.003, compared to autoregressive LM rescoring (4.19% WER, RTF 0.008) and first-pass AR decoding (3.86% WER, RTF 0.078). MDLM outperforms USDM in static rescoring on this data scale, whereas USDM excels in joint decoding due to its full-vocabulary distribution.
| System | Decoding Method | dev-other WER [%] | RTF |
|---|---|---|---|
| CTC Baseline | Greedy | 5.08 | 0.002 |
| AR LM (24L, 1024 dim) | First-pass | 3.86 | 0.078 |
| MDLM (24L, 25 ep, K=256) | Rescoring | 4.47 | 2.040 |
| USDM (24L, 10 ep, K=256) | Rescoring | 4.71 | 2.086 |
| CTC-USDM (24L, L=1, t_start=0.1) | Joint-Decoding | 4.66 | 0.003 |
Limitations
The study evaluates diffusion LMs exclusively on the LibriSpeech corpus, leaving multilingual and low-resource generalizability unverified. Diffusion LM rescoring remains computationally expensive at high Monte Carlo sample counts (e.g., K = 256 yields an RTF around 2.04). Furthermore, autoregressive LMs still outperform diffusion-based approaches in raw accuracy on both rescoring and first-pass tasks under the tested data scale.
Why read this
Speech and ML researchers looking for non-autoregressive alternatives to standard autoregressive LMs for ASR will find a rigorous formulation of MDLM rescoring adjustments and a pioneering blueprint for token-level joint decoding between CTC and uniform-state diffusion models.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Automatic speech recognition, high-throughput speech transcription pipelines, and non-autoregressive speech-to-text decoding.
Institutions
RWTH Aachen University, AppTek
Funding / 經費: NeuroSys, Federal Ministry of Research, Technology and Space BMFTR, RESCALE, Federal Ministry for the Environment, Nature Conservation, Nuclear Safety and Consumer Protection, Federal Ministry of Education and Research
Related
- MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding — same problem · relatedness 2.9/3
- Refining the Latent Bridge: Superior ASR Performance via Adapter-Only Alignment with Diffusion LLMs — same problem · relatedness 2.8/3
- LLM-as-Joiner: Decoupling Alignment from Language Modeling in Label-synchronous ASR — same problem · relatedness 2.0/3
- Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy — same problem · relatedness 2.0/3
- Controllable Accent Normalization via Discrete Diffusion — shared technique · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2070