All papers
Speech LLMs & dialogueFull-paper digest

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Haoyu Zhang, Jiaxian Guo, Dong Yang, Yusuke Iwasawa, Yutaka Matsuo

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.3 KB · Ready to paste

Preview copied content

TL;DR — AQA-TTRL is a test-time reinforcement learning framework that enables Large Audio Language Models to autonomously self-adapt on unlabeled test data, yielding average accuracy gains of 4.42% for 7B models and 11.04% for 3B models.

Key contributions

  • Proposes AQA-TTRL, a label-free test-time adaptation framework for audio question answering using Group Relative Policy Optimization (GRPO) driven by self-generated pseudo-labels.
  • Introduces a confidence-weighted advantage scheme that re-scales training gradients based on majority-vote consistency, prioritizing high-reliability pseudo-labels.
  • Develops a multiple-attempt sampling strategy to combat rollout collapse and vanishing policy advantages without requiring memory-prohibitive group sizes.
  • Provides empirical validation across MMAU, MMAR, and MMSU benchmarks, showing that an adapted 3B model can outperform a static 7B model.

Problem

Large Audio Language Models (LALMs) suffer from acoustic mismatch and domain shifts caused by background noise, recording variations, and speaker diversity in real-world deployments. While collecting human-annotated data for supervised fine-tuning resolves this, it is prohibitively expensive and time-consuming. Prior approaches cannot adapt on-the-fly without gold labels, making label-free test-time adaptation essential. However, applying Test-Time Reinforcement Learning (TTRL) to audio processing is challenging due to inherent noise in self-generated pseudo-labels and the "advantage collapse" phenomenon where high-confidence responses yield zero advantage.

Method

AQA-TTRL operates in two stages: pseudo-label generation via majority voting and pseudo-label guided policy updates. For each audio-question pair (a,q)(a, q), the model generates MM stochastic predictions (T=1T=1), and majority voting yields a consensus pseudo-label y^\hat{y}. The pseudo-label confidence is quantified as the fraction of votes matching y^\hat{y}.

For policy optimization, the framework utilizes Group Relative Policy Optimization (GRPO) with a group size G=4G=4, using binary exact-match rewards combining format and accuracy (ri=racc(oi,y^)+rformat(oi)r_i = r_{acc}(o_i, \hat{y}) + r_{format}(o_i)). To address noisy pseudo-labels and rollout collapse, two mechanisms are introduced. First, a confidence-weighted advantage scales normalized advantages using an exponential function f(Conf)=exp⁡(Conf)f(Conf) = \exp(Conf) to prioritize reliable signals while bounding amplification. Second, a multiple-attempt sampling strategy sequentially draws groups of responses (G1,G2,G3G_1, G_2, G_3) and selects the first group containing non-identical outputs to bypass reward stagnation.

Training uses AdamW with a learning rate of 1e−61e-6, weight decay of 0.010.01, global batch size of 8 across 4 GPUs, gradient clipping of 1.0, and bf16 precision. Hyperparameters ϵ=0.2\epsilon=0.2 and β=0\beta=0, with updates running for 100 steps on smaller datasets and 500 steps on larger datasets.

Experimental setup

Evaluated on MMAU (test-mini and test), MMAR, and MMSU benchmarks using Qwen2.5-Omni 7B and 3B as base models. Compared against Direct Inference (DI), Direct Inference with Majority Voting (DIMV), and Supervised Fine-Tuning (SFT) on identical pseudo-labels for 3 epochs. Metrics include accuracy percentages across sound, music, and speech subsets.

Results

On Qwen2.5-Omni 7B, AQA-TTRL improves average accuracy from 64.39% (DI) and 65.59% (DIMV) to 68.81%, with notable gains on MMAU test-mini (76.80%) and MMAR (63.20%). On Qwen2.5-Omni 3B, average accuracy jumps from 53.82% (DI) to 64.86%, allowing the adapted 3B model to outperform the unadapted 7B model (64.39%). Ablation studies confirm that combining confidence weighting and multiple-attempt sampling yields the highest synergy, consistently outperforming standalone G-MV (67.50% average).

SystemMMAU test-miniMMAU testMMARMMSUAverage
Qwen2.5-Omni 7B (DI)72.4070.6057.7056.8464.39
Qwen2.5-Omni 7B (DIMV)73.3072.0158.9058.1665.59
Qwen2.5-Omni 7B (SFT)73.9071.7458.6058.0265.57
Qwen2.5-Omni 7B (Ours)76.8073.7463.2061.4868.81
Qwen2.5-Omni 3B (DI)61.5061.5546.9045.3453.82
Qwen2.5-Omni 3B (Ours)72.3071.0557.9058.1864.86

Limitations

The framework assumes tasks can be framed as closed-form or easily verified answer choices using exact-match rewards, limiting direct application to open-ended speech generation or conversational synthesis. The approach relies on multi-sample rollout generation during test time, which introduces computational overhead compared to single-pass inference. Evaluation is restricted to English-centric or standard public audio question-answering benchmarks, leaving multilingual or streaming long-form audio scenarios untested.

Why read this

Researchers and engineers working on test-time adaptation, reinforcement learning without ground truth labels, or deploying resource-efficient audio language models will find this a blueprint for bypassing costly supervised fine-tuning.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

On-the-fly acoustic adaptation for smart speakers, edge voice assistants, and audio surveillance systems operating in changing acoustic environments without manual data relabeling.

Institutions

University of Tokyo

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-288