TL;DR — This paper introduces attention-guided reliability scaling for contrastive decoding in LLM-based audio-visual speech recognition (AVSR), achieving consistent Word Error Rate reductions across clean and severely noisy environments without any model fine-tuning.
Key contributions
- Formulates a training-free, inference-time contrastive decoding (CD) framework for AVSR that contrasts audio-visual conditioning (Expert) against audio-only conditioning (Amateur) within the same LLM.
- Analyzes the fundamental noise-robustness vs. clean-speech preservation trade-off inherent in static fixed-weight contrastive decoding.
- Proposes a multiplicative soft-gating mechanism driven by relative audio energy, audio attention entropy, and Jensen-Shannon predictive divergence to dynamically modulate token-level intervention strength.
- Demonstrates consistent generalization across different model scales (0.5B to 8B parameters) and out-of-distribution evaluation sets (LRS2).
Problem
Large language model-based AVSR systems are vulnerable to degraded acoustic inputs because they can over-rely on noisy audio when environmental conditions deteriorate. Prior strategies to fix this require structural modifications, auxiliary gating modules, or additional fine-tuning, which increases training cost and deployment complexity. Applying standard contrastive decoding with a static weight helps under severe noise but over-corrects and distorts reliable predictions during clean conditions. Because acoustic signal-to-noise ratios fluctuate dynamically at the token level, a uniform intervention strength fails to balance noise robustness with clean-speech preservation.
Method
The framework operates entirely at inference time by running the underlying LLM-based AVSR model under two conditioning modes: full audio-visual input for the Expert, and audio-only input (dropping video embeddings) for the Amateur. The standard contrastive weight is modulated per token via a multiplicative soft gate , where smoothly interpolates between standard AVSR () and full contrastive decoding ().
The relative audio energy measures how much attention the final decoding token assigns to the audio region at the last Transformer layer, averaged across all attention heads. To handle volume and SNR fluctuations, is dynamically normalized against a running utterance mean through a sigmoid function with sensitivity parameter . The audio entropy is calculated independently per attention head over normalized audio weights, divided by to bound it in , and passed through a sigmoid with to measure acoustic uncertainty.
The Jensen-Shannon divergence measures predictive disagreement between Expert and Amateur distributions, normalized by . To prevent rank distortion—where extreme Amateur collapse injects massive negative log-probability offsets across the vocabulary—a Gaussian filter centered at with suppresses intervention when distributions are either identical or excessively divergent, focusing contrast on informative transitional states. The base contrastive weight is set to , and gating parameters use .
Experimental setup
Evaluated on the LRS3 dataset (433 hours training set, 1,327 test utterances) and out-of-distribution on the LRS2 test set. Noise from the MUSAN dataset (equal mix of noise, speech, and music) was artificially injected at clean, 0 dB, -5 dB, -10 dB, and -15 dB SNR levels. Three model scales were tested: Llama-AVSR (Llama-3.1-8B with Whisper-Medium), Omni-AVSR (Llama-3.2-1B with Whisper-Small), and Qwen-AVSR (Qwen2.5-0.5B with Whisper-Small), using frozen AV-HuBERT-Large visual encoders and LoRA projection layers. Experiments ran on NVIDIA A100 GPUs, adding an 8.6% per-utterance latency overhead (+136.4 ms over a baseline of 1577.9 ms).
Results
On the Llama-AVSR 8B LRS3 test set, the proposed method improves clean WER from 0.0095 to 0.0082 (+13.68% relative) and severely noisy -15 dB WER from 0.3723 to 0.3369 (+9.51% relative), achieving an average relative improvement of 9.95% across all SNR conditions. Across Llama-AVSR, Omni-AVSR, and Qwen-AVSR on LRS3 and LRS2, the method consistently boosts performance in both clean and degraded conditions, outperforming static fixed- contrastive decoding which forces a trade-off between clean accuracy and denoising. Ablation studies confirm that combining energy, entropy, and JS divergence cues yields superior balance across all noise tiers compared to using any single cue in isolation.
| System & Condition | Clean (WER) | 0 dB (WER) | -5 dB (WER) | -10 dB (WER) | -15 dB (WER) |
|---|---|---|---|---|---|
| Llama-AVSR 8B (Baseline) | 0.0095 | 0.0367 | 0.0945 | 0.2358 | 0.3723 |
| Llama-AVSR 8B (Ours) | 0.0082 | 0.0330 | 0.0866 | 0.2167 | 0.3369 |
| Omni-AVSR 1B (Baseline) | 0.0142 | 0.0495 | 0.1175 | 0.2509 | 0.3305 |
| Omni-AVSR 1B (Ours) | 0.0111 | 0.0468 | 0.1108 | 0.2306 | 0.3053 |
Limitations
The evaluation relies on synthetic noise injections via MUSAN rather than natural multi-condition acoustic recordings with reverberation and spatial distortion. The approach introduces a small inference latency overhead (~8.6%) due to evaluating both audio-visual and audio-only conditioning passes at each step. Additionally, hyperparameter settings like and sigmoid sharpness rely on validation tuning which may require re-calibration for entirely different backbone architectures or low-resource languages.
Why read this
Speech and ML researchers working on LLM-based speech recognition or multi-modal fusion will learn how to stabilize inference-time generation against modality corruption without retraining. It provides a blueprint for leveraging attention dynamics and predictive divergence to dynamically regulate contrastive decoding.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Robust audio-visual speech recognition systems for edge devices, noisy automotive environments, and multi-modal meeting transcription.
Institutions
Hanyang University
Funding / 經費: Institute of Information & Communications Technology Planning & Evaluation, Korea government (MSIT), Artificial Intelligence Graduate School Program (Hanyang University)
Related
- VIB-AVSR: Variational Information Bottleneck for Noise-Robust LLM-Based Audio-Visual Speech Recognition — same problem · relatedness 2.8/3
- Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement — same problem · relatedness 2.7/3
- Adaptive AVSR: Integrating Speaker and Environmental Embeddings for Robust Audio-Visual Speech Recognition — same problem · relatedness 2.4/3
- Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding — shared technique · relatedness 2.2/3
- Training-Free Intelligibility-Guided Observation Addition for Noisy ASR — same problem · relatedness 2.2/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-929