All papers
Audio understandingFull-paper digest

TAD: Token-Adaptive Contrastive Decoding with Confidence-Guided Gating for Hallucination Mitigation in Large Audio-Language Models

Heyu Chang, Nianwen Si, Hao Zhang, Wenlin Zhang, Dan Qu

Code & resourcesgithub.com/Changhy26/TAD

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — Token-Adaptive Decoding (TAD) is a training-free inference strategy that suppresses audio object hallucinations in large audio-language models by applying a confidence-guided penalty to affirmative tokens at the first decoding step. It improves F1 by up to 0.117 on AudioCaps-Hallucination compared to fixed-strength contrastive baselines.

Key contributions

  • Proposes a decision-aware, training-free plug-in logits processor called Token-Adaptive Decoding (TAD) for large audio-language models.
  • Employs a robust subword union strategy with log-sum-exp pooling over surface variants to aggregate vocabulary-level logits into semantic YES/NO categories.
  • Introduces a confidence-margin-guided gate applied exclusively at the first decoding step based on the margin change between real audio and silent references.
  • Demonstrates consistent F1 and recall improvements over Audio-Aware Decoding (AAD) and Adaptive Vector Steering (AVS) on AudioCaps-Hallucination and Clotho-AQA.

Problem

Large audio-language models (LALMs) frequently generate audio object hallucinations, answering 'yes' to binary presence questions even when the target sound is absent, particularly for frequent categories and adversarial prompts. Prior training-free methods like Audio-Aware Decoding (AAD) and Adaptive Vector Steering (AVS) rely on global, fixed-strength contrastive reweighting across all decoding steps without adapting to varying audio evidence strength. Because binary audio question answering (AQA) decisions are predominantly locked in at the first decoding step, failing to condition interventions on initial confidence leads to overcorrection or insufficient hallucination suppression.

Method

TAD operates by contrasting logits produced under real audio conditioning with those under a matched silent reference (an all-zero waveform of identical length), combined via an update weight α\alpha (0.50.5 or 1.01.0). To make the vocabulary projection robust against subword tokenization artifacts, the method enumerates surface variants for 'yes' and 'no' (e.g., capitalized, spaced, and lowercase forms) and applies log-sum-exp pooling over these subword sets to extract scalar preferences utyesu_t^{\text{yes}} and utnou_t^{\text{no}} for both the audio and silent branches.

At the first decoding step (t=1t=1), TAD computes the YES/NO margin under real audio (maudiom^{\text{audio}}) and under silence (msilentm^{\text{silent}}), yielding an audio-induced margin gain δ=maudio−msilent\delta = m^{\text{audio}} - m^{\text{silent}}. A threshold τ\tau (set to 0.20.2) checks if the audio evidence is weak; if δ<τ\delta < \tau, an additive penalty γ=2.5\gamma = 2.5 is applied exclusively to affirmative tokens in the contrastive logits. For all subsequent decoding steps (t>1t > 1), standard AAD logits are used without gating, ensuring the intervention targets only the initial binary decision boundary.

Experimental setup

Evaluated on AudioCaps-Hallucination (test split split into RANDOM with 30,220 pairs, ADVERSARIAL with 31,047 pairs, and POPULAR with 31,376 pairs) and a binary yes/no subset of Clotho-AQA containing 1,991 samples. Compared against default decoding, AVS, and AAD with contrast weights α∈{0.5,1.0}\alpha \in \{0.5, 1.0\}. Evaluated using Accuracy, Precision, Recall, and F1 (treating 'no' as the positive class). Implemented on Qwen2-Audio-7B-Instruct and Gemma-3n-E4B-it.

Results

On AudioCaps-Hallucination using Qwen2-Audio-7B-Instruct (Random split), TAD achieves an F1 of 0.853 (vs 0.736 for AAD and 0.805 for AVS) and improves F1 by 0.059 to 0.117 across Random, Adversarial, and Popular splits compared to AAD. On Clotho-AQA with Qwen2, TAD attains the highest F1 of 0.816 (with α=1.0\alpha=1.0) and a recall of 0.907, outperforming AAD's F1 of 0.810. However, on the smaller Gemma-3n-E4B-it model, while recall increases substantially, precision drops due to a conservative bias, indicating the intervention can occasionally overcorrect weaker backbones.

SystemSettingAccuracyPrecisionRecallF1
Default (Qwen2)Random0.5930.7690.2660.395
AAD [11] (α=1.0\alpha=1.0)Random0.7620.8240.6660.736
AVS [20]Random0.7730.7060.9370.805
TAD (ours, α=1.0\alpha=1.0)Random0.8410.8470.8580.853
TAD (ours, α=1.0\alpha=1.0)Clotho-AQA0.7960.7410.9070.816

Limitations

Evaluated exclusively on binary AQA datasets with yes/no questions, leaving open-ended hallucination tasks unexplored. The fixed hyperparameters (τ=0.2\tau = 0.2, γ=2.5\gamma = 2.5) were not exhaustively tuned per architecture, leading to excessive conservativeness and precision drops on smaller multimodal models like Gemma-3n-E4B-it. The method assumes the binary decision is entirely formed at the first token, which may not generalize to multi-word or multi-token decision formats.

Why read this

Speech and ML engineers building deployment-ready audio question answering systems will find a practical, training-free plug-and-play decoding algorithm that drastically reduces affirmative hallucinations. Researchers will appreciate the rigorous token-set log-sum-exp pooling and first-step confidence margin analysis.

Code

Applications

Audio question answering systems, acoustic event detection validators, and reliable voice-controlled multimodal assistants.

Institutions

Information Engineering University

Funding / 經費: Henan Province Major Industrial “Challenge-Based Innovation”, Natural Science Foundation of Henan, Science and Technology Key Project Plan of Henan

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-637