TL;DR — AVUR-LLM is an audio-visual speech recognition framework that uses sparse cross-modal alignment, confidence-aware decoding, and discrete visual unit-guided LLM rescoring, achieving a 37% relative WER reduction over baselines at 0 dB SNR.
Key contributions
- Sparse Modality Alignment (SMA): inserts lightweight cross-modal attention blocks into upper audio encoder layers with stop-gradients to preserve pristine acoustic representations.
- Adaptive Modulated Fusion (AMF): computes token-level acoustic entropy uncertainty to dynamically scale and gate visual feature injection during Whisper decoding.
- Visual Unit-Guided Refinement (VUR): discretizes mid-layer AV-HuBERT features via a 2000-cluster K-means codebook and run-length compression to prompt a LoRA-adapted LLaMA-2 7B for N-best hypothesis rescoring.
- A two-stage training scheme that decouples first-stage multimodal acoustic-visual decoding from second-stage LLM-based sequence refinement.
Problem
Prior LLM-based AVSR systems either project continuous audio and visual features directly into the LLM or apply shallow feature concatenation, which tightly couples the modalities, increases memory footprint, over-sensitizes the model to input noise, and lacks fine-grained control. Additionally, naive multimodal LLM prompts suffer from high computational overhead and fail to effectively balance acoustic degradation under adverse noise conditions. This matters because robust speech recognition in real-world acoustic environments requires stable, controlled cross-modal fusion that can dynamically lean on visual cues when audio channels degrade.
Method
The framework operates in two distinct stages. In Stage 1, acoustic features are extracted using Whisper Medium (80-dimensional log-Mel filterbanks) and visual features via AV-HuBERT Large (from 96x96 cropped lip regions at 25 fps). The Sparse Modality Alignment (SMA) module upsamples visual features via a differentiable resampler and maps them to the audio feature dimension, then injects three cross-attention blocks into the upper layers of the Whisper encoder where visual features act as queries and audio features act as keys/values protected by stop-gradients. Simultaneously, an Adaptive Modulated Fusion (AMF) module inside the Whisper decoder calculates a forward-only acoustic probe at each layer, deriving token-wise acoustic uncertainty entropy to govern learnable sigmoid amplitude gates and tanh direction gates that regulate cross-attention and feed-forward residual injection.
In Stage 2, the intermediate visual representations from the 12th layer of AV-HuBERT are discretized using an offline K-means codebook (size K=2000). Run-length compression is applied to collapse maximal contiguous spans of identical token indices into single averaged tokens. These discrete visual tokens form an explicit prompt alongside an N-best candidate list generated by Stage 1, feeding a LLaMA-2 7B model fine-tuned via LoRA (rank 16, dropout 0.05 on query, key, value, and output projection layers). The LLM is optimized using a list-wise softmax loss over the candidate scores to perform robust sequence-level rescoring.
Experimental setup
Evaluated on the LRS3 dataset (433 hours of TED talks, plus a 30-hour trainval subset and an extended 1759-hour setting combining LRS3 with VoxCeleb2). Noise robustness is tested by adding MUSAN babble noise at SNRs of 10, 5, 0, -5, and -10 dB. Performance is measured using Word Error Rate (WER). Implemented using Whisper Medium, AV-HuBERT Large, and LLaMA-2 7B, trained with the AdamW optimizer (learning rate 1e-4 for Stage 1, 5e-4 for Stage 2; weight decay 0.1).
Results
Under clean conditions on the 433h LRS3 split, AVUR-LLM achieves an audio-visual WER of 0.75%, outperforming MMS-LLaMA (0.9%) and Whisper-Flamingo (1.0%). At the expanded 1759-hour data scale, the AVSR WER drops further to 0.68%. Under adverse noisy conditions, the model demonstrates pronounced robustness, securing a WER of 1.7% at 0 dB SNR (a 37% relative reduction over MMS-LLaMA's 2.7%) and maintaining 6.3% WER at -5 dB SNR.
Ablation experiments reveal that removing VUR degrades clean WER to 0.97% and 0 dB WER to 5.50%, while dropping AMF (SMA only) causes a sharp collapse to 1.30% clean and 6.70% noisy, proving that confidence-aware fusion and visual refinement drive noise resilience. Feature extraction depth analysis shows that the 12th AV-HuBERT layer with a 2K codebook size outperforms both shallow (layer 3) and deep (layer 18) alternatives.
| System | Clean | 10 dB | 5 dB | 0 dB | -5 dB | -10 dB |
|---|---|---|---|---|---|---|
| AV-HuBERT [7] | 1.4 | 2.0 | 2.6 | 5.8 | 16.6 | 34.9 |
| CMA [34] | 1.5 | 1.8 | 2.4 | 4.4 | 11.9 | 25.8 |
| Whisper-Flamingo [13] | 1.0 | 1.4 | 2.0 | 6.3 | 28.0 | 42.6 |
| Llama-AVSR [16] | 0.95 | — | 2.2 | 4.2 | 16.9 | — |
| MMS-LLaMA [17] | 0.9 | — | 1.3 | 2.7 | 7.4 | — |
| AVUR-LLM (Full) | 0.75 | 1.0 | 1.4 | 1.7 | 6.3 | 12.9 |
Limitations
Evaluated exclusively on English-language datasets (LRS3 and VoxCeleb2), leaving multilingual generalization untested. The visual front-end relies on clean face-cropped bounding boxes via Dlib face tracking, which may degrade under unconstrained, real-world head poses or extreme illumination changes. The two-stage training pipeline also decouples first-stage decoding from second-stage LLM rescoring, preventing end-to-end gradient propagation through the language model.
Why read this
Researchers building multi-modal speech LLMs will find valuable design patterns here, specifically how to avoid destabilizing pretrained acoustic encoders via sparse upper-layer alignment and how to stabilize inference using entropy-gated fusion paired with compact discrete visual prompting.
Code
Applications
Robust automatic speech recognition for noisy environments, video conferencing enhancement, hearing assistive devices, and audio-visual transcription systems.
Institutions
Wuhan University, Chinese University of Hong Kong, Shenzhen, OPPO
Funding / 經費: National Natural Science Foundation of China, Yangtze River Delta Science and Technology Innovation Community Joint Research Project, OPPO
Related
- VIB-AVSR: Variational Information Bottleneck for Noise-Robust LLM-Based Audio-Visual Speech Recognition — same problem · relatedness 2.8/3
- Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition — same problem · relatedness 2.7/3
- Adaptive AVSR: Integrating Speaker and Environmental Embeddings for Robust Audio-Visual Speech Recognition — same problem · relatedness 2.6/3
- Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition — same problem · relatedness 2.0/3
- DASH: Dual-View Self-Distillation with Multi-Layer Hidden Representations for Robust Speech Recognition — same problem · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1277