All papers
Resources & evaluationFull-paper digest

Hearing the Order: Investigating Position Bias in Large Audio-Language Models

Yu-Xiang Lin, Chen-An Li, Sheng-Lun Wei, Po-Chun Chen, Hsin-Hsi Chen, Hung-yi Lee

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.2 KB · Ready to paste

Preview copied content

TL;DR — This paper provides the first systematic study of position bias in Large Audio-Language Models (LALMs) on multiple-choice benchmarks, demonstrating that shuffling answer options causes accuracy fluctuations of up to 24% and alters model leaderboards. The authors show that permutation-based mitigation strategies effectively stabilize evaluation outcomes.

Key contributions

  • First comprehensive investigation of position bias in LALMs across six models, three text/speech benchmarks, and their spoken counterparts.
  • Construction of spoken benchmark variants (SPEECH-MMAU, SPEECH-MMAR, SPEECH-MMLU) via GPT-4o mini TTS to analyze modality effects.
  • Demonstration that option identifiers (A, B, C, D) improve raw accuracy but fail to mitigate underlying positional preferences.
  • Evaluation of cyclic and full-permutation decoding strategies as robust mitigation techniques that suppress variance and rectify distorted model rankings.

Problem

Multiple-choice benchmarks are increasingly used to evaluate the reasoning capabilities of Large Audio-Language Models, but they risk conflating true comprehension with structural position bias. While text LLMs and vision-language models have been widely scrutinized for positional sensitivity, this phenomenon remains unexplored in the audio-language domain. Because model predictions can swing heavily depending on whether the correct answer appears first or last, current standardized evaluations may yield misleading performance profiles and flawed model rankings.

Method

The study evaluates six representative LALMs varying in architecture and scale: Gemini-2.0-Flash, Phi-4-Multimodal, Qwen2.5-Omni (3B and 7B), and Voxtral-Mini-3B and Voxtral-Small-24B. Experiments are executed using the OpenAI simple-eval protocol with a temperature of 0 and a maximum generation length of 1024 tokens. Spoken benchmarks are generated by converting text prompts to audio using GPT-4o mini TTS, filtering out audio sequences exceeding 180 seconds to maintain computational tractability.

To diagnose and mitigate bias, the authors systematically reassign the correct ground-truth label to positions A, B, C, and D while randomly shuffling distractors. Furthermore, they implement cyclic and full-permutation evaluation protocols. In these strategies, each permuted choice ordering is fed as an independent input, and the final prediction is determined via majority voting across permutations, functioning analogously to test-time self-consistency.

Experimental setup

Evaluations are conducted on three core benchmarks: MMAU (test-mini subset, 933 samples), MMAR (815 samples), and MMLU (14,019 samples), alongside their spoken counterparts (SPEECH-MMAU, SPEECH-MMAR, SPEECH-MMLU). Performance is measured using Accuracy, Delta Accuracy, Relative Standard Deviation (RSD), and Choice Kullback-Leibler Divergence (CKLD). Models include 3B to 24B parameter architectures evaluated at zero temperature.

Results

All six evaluated LALMs exhibit pronounced position bias, with accuracy fluctuations reaching up to 24% for Phi-4-Multimodal when the correct answer position is manipulated. Applying full-permutation mitigation consistently suppresses bias metrics and boosts accuracy; for instance, Qwen2.5-Omni-7B rises from an original MMAU accuracy of 73.96% to 78.03% under full permutation. Furthermore, shifting option orders directly alters model rankings on leaderboards, demonstrating that standard single-run evaluations fail to reflect true model capabilities.

System / ConditionMMAU AccuracySPEECH-MMAU AccuracyMMAR AccuracySPEECH-MMAR Accuracy
Phi-4-Multimodal (Original)65.27%53.06%43.19%41.47%
Phi-4-Multimodal (Full Permutation)69.24%58.31%51.53%46.50%
Qwen2.5-Omni-7B (Original)73.96%71.92%53.25%52.27%
Qwen2.5-Omni-7B (Full Permutation)78.03%75.46%56.69%56.07%
Gemini-2.0-Flash (Original)74.17%73.74%64.54%63.07%
Gemini-2.0-Flash (Full Permutation)75.67%75.56%67.97%65.52%

Limitations

The study is bounded by filtering out audio samples longer than 180 seconds and restricting evaluations to four-option multiple-choice questions. Permutation-based mitigation strategies introduce significant computational overhead at inference time, multiplying API or compute costs by the number of permutations. Additionally, synthetic TTS audio generation may not fully capture the acoustic diversity of natural human speech corpora.

Why read this

Researchers and benchmark designers evaluating LALMs should read this paper to understand how severely option ordering distorts model rankings and to adopt permutation-based evaluation protocols for trustworthy results.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Standardized evaluation frameworks for audio-language models, speech-based automated testing systems, and robust multi-modal benchmark design.

Institutions

National Taiwan University

Funding / 經費: Ministry of Education, Taiwan Centers of Excellence in Artificial Intelligence, NTU Artificial Intelligence Center of Research Excellence

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1025