TL;DR — This paper introduces Online Predictive Coding (OPC) and dual-mode Layer Normalization to bridge the performance gap between streaming and non-streaming speech models, cutting word error rates on LibriSpeech test-clean at 160 ms latency from 3.65% to 3.40%.
Key contributions
- Proposed Online Predictive Coding (OPC), which regularizes learnable online registers via multi-step future feature prediction during self-supervised pre-training.
- Adopted Dual-mode Layer Normalization, maintaining separate affine parameters ( and ) for online and offline pathways to prevent mode-specific distribution shifts.
- Demonstrated consistent WER improvements across LibriSpeech (test-clean and test-other) and WSJ benchmarks without increasing algorithmic latency.
Problem
Leading self-supervised speech models like wav2vec 2.0, HuBERT, and BEST-RQ are trained exclusively in offline (non-streaming) modes, causing performance drops when deployed for real-time online streaming where future context is absent. Prior dual-mode frameworks that share encoder parameters across streaming and non-streaming modes suffer from optimization instability and attention mismatches. Although past work introduced learnable 'online registers' to mimic missing future context, they lacked explicit supervision to robustly encode future information, leaving performance gains marginal.
Method
The model builds on a wav2vec 2.0 BASE architecture featuring a convolutional feature encoder and a Transformer backbone. For online streaming mode, speech features are partitioned into chunks of size (varied via Dynamic Chunk Training from 2 to 32) with optional lookahead . A single learnable online register () is appended to each chunk to act as a surrogate for unseen future frames, enabling parallel chunk-based attention masking during training.
To ensure these registers capture future context, Online Predictive Coding (OPC) is introduced. Using linear projections , the representations of the online registers jointly predict future offline representations extracted from the non-streaming pathway. The OPC objective () minimizes the cosine distance between the predicted vectors and the unmasked offline targets, with stop-gradient applied to the offline targets to prevent collapse. This loss is jointly optimized alongside the standard wav2vec 2.0 online and offline masked prediction losses (, ) and codebook diversity loss () using weight hyperparameters and .
To stabilize parameter sharing across modes—particularly given that the active online registers introduce activation patterns distinct from standard speech frames—Dual-mode Layer Normalization is implemented. Every LayerNorm layer uses decoupled affine parameters () for online and offline passes while sharing all other weights. Pre-training uses 16 NVIDIA H200 GPUs for 100k steps on 960 hours of unlabeled LibriSpeech data, using the Adam optimizer with a learning rate warming up to over 8k steps.
Experimental setup
Pre-trained on the 960-hour LibriSpeech corpus. Fine-tuned on LibriSpeech 960h and WSJ (train_si284) using Connectionist Temporal Classification (CTC) loss for 320k steps on 8 NVIDIA H200 GPUs. Baselines include offline wav2vec 2.0, standard dual-mode baseline without registers, dual-mode with registers, and UFO2. Evaluated using Word Error Rate (WER %) on LibriSpeech (test-clean, test-other) and WSJ (eval92, eval93) using a 4-gram language model via Flashlight beam search.
Results
On LibriSpeech at a low-latency 160 ms setting (), adding online registers drops online WER from 3.65% to 3.50% on test-clean and 10.15% to 9.80% on test-other. Adding OPC further drops online WER to 3.40% on test-clean and 9.65% on test-other, while simultaneously improving offline WER from 2.73% to 2.64% (test-clean) and 6.63% to 6.41% (test-other). Ablations on the number of predicted future frames () show that performance drops if is too small () or too large (), peaking at . Where it does not win: on WSJ eval93, OPC shows a slight 1.2% relative offline WER increase compared to the baseline, indicating mild domain bias from the auxiliary pre-training task.
| System / Condition | test-clean (Offline) | test-clean (Online) | test-other (Offline) | test-other (Online) |
|---|---|---|---|---|
| Dual-mode Only | 2.73% | 3.65% | 6.63% | 10.15% |
| w/ Online Registers | 2.70% | 3.50% | 6.52% | 9.80% |
| w/ Online Predictive Coding (Ours) | 2.64% | 3.40% | 6.41% | 9.65% |
| wav2vec 2.0 (Offline-only baseline) | 2.6% | - | 6.1% | - |
| UFO2 | 3.0% | 3.8% | 7.1% | 9.4% |
| Ours (640 ms chunk size, ) | 2.6% | 3.1% | 6.4% | 8.3% |
Limitations
Evaluated exclusively on English corpora (LibriSpeech and WSJ) and restricted to Automatic Speech Recognition downstream evaluation. The auxiliary OPC task exhibits minor cross-domain performance degradation when transferring to out-of-domain evaluation sets like WSJ. The compute requirements demand large-scale resources (16x H200 GPUs for pre-training).
Why read this
Speech researchers building real-time streaming speech recognition models will find this a practical guide to eliminating the performance penalty of streaming self-supervised models without architectural bloat. It provides clear recipes for combining predictive coding objectives with dual-mode normalization layers.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Real-time streaming automatic speech recognition, voice assistants, and on-device live captioning.
Institutions
LY Corporation, Carnegie Mellon University
Related
- Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization — same problem · relatedness 2.5/3
- Improving streaming ASR with foundation models using emission policies — same problem · relatedness 2.1/3
- BACON: Boundary-Aware Convolution for Streaming Conformer Models — same problem · relatedness 2.0/3
- TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling — same problem · relatedness 2.0/3
- Token-Independent Language Representations for Low-Latency Configurable Multilingual Speech Recognition — same problem · relatedness 1.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1997