TL;DR — This paper provides the first systematic study of position bias in Large Audio-Language Models (LALMs) on multiple-choice benchmarks, demonstrating that shuffling answer options causes accuracy fluctuations of up to 24% and alters model leaderboards. The authors show that permutation-based mitigation strategies effectively stabilize evaluation outcomes.
Key contributions
- First comprehensive investigation of position bias in LALMs across six models, three text/speech benchmarks, and their spoken counterparts.
- Construction of spoken benchmark variants (SPEECH-MMAU, SPEECH-MMAR, SPEECH-MMLU) via GPT-4o mini TTS to analyze modality effects.
- Demonstration that option identifiers (A, B, C, D) improve raw accuracy but fail to mitigate underlying positional preferences.
- Evaluation of cyclic and full-permutation decoding strategies as robust mitigation techniques that suppress variance and rectify distorted model rankings.
Problem
Multiple-choice benchmarks are increasingly used to evaluate the reasoning capabilities of Large Audio-Language Models, but they risk conflating true comprehension with structural position bias. While text LLMs and vision-language models have been widely scrutinized for positional sensitivity, this phenomenon remains unexplored in the audio-language domain. Because model predictions can swing heavily depending on whether the correct answer appears first or last, current standardized evaluations may yield misleading performance profiles and flawed model rankings.
Method
The study evaluates six representative LALMs varying in architecture and scale: Gemini-2.0-Flash, Phi-4-Multimodal, Qwen2.5-Omni (3B and 7B), and Voxtral-Mini-3B and Voxtral-Small-24B. Experiments are executed using the OpenAI simple-eval protocol with a temperature of 0 and a maximum generation length of 1024 tokens. Spoken benchmarks are generated by converting text prompts to audio using GPT-4o mini TTS, filtering out audio sequences exceeding 180 seconds to maintain computational tractability.
To diagnose and mitigate bias, the authors systematically reassign the correct ground-truth label to positions A, B, C, and D while randomly shuffling distractors. Furthermore, they implement cyclic and full-permutation evaluation protocols. In these strategies, each permuted choice ordering is fed as an independent input, and the final prediction is determined via majority voting across permutations, functioning analogously to test-time self-consistency.
Experimental setup
Evaluations are conducted on three core benchmarks: MMAU (test-mini subset, 933 samples), MMAR (815 samples), and MMLU (14,019 samples), alongside their spoken counterparts (SPEECH-MMAU, SPEECH-MMAR, SPEECH-MMLU). Performance is measured using Accuracy, Delta Accuracy, Relative Standard Deviation (RSD), and Choice Kullback-Leibler Divergence (CKLD). Models include 3B to 24B parameter architectures evaluated at zero temperature.
Results
All six evaluated LALMs exhibit pronounced position bias, with accuracy fluctuations reaching up to 24% for Phi-4-Multimodal when the correct answer position is manipulated. Applying full-permutation mitigation consistently suppresses bias metrics and boosts accuracy; for instance, Qwen2.5-Omni-7B rises from an original MMAU accuracy of 73.96% to 78.03% under full permutation. Furthermore, shifting option orders directly alters model rankings on leaderboards, demonstrating that standard single-run evaluations fail to reflect true model capabilities.
| System / Condition | MMAU Accuracy | SPEECH-MMAU Accuracy | MMAR Accuracy | SPEECH-MMAR Accuracy |
|---|---|---|---|---|
| Phi-4-Multimodal (Original) | 65.27% | 53.06% | 43.19% | 41.47% |
| Phi-4-Multimodal (Full Permutation) | 69.24% | 58.31% | 51.53% | 46.50% |
| Qwen2.5-Omni-7B (Original) | 73.96% | 71.92% | 53.25% | 52.27% |
| Qwen2.5-Omni-7B (Full Permutation) | 78.03% | 75.46% | 56.69% | 56.07% |
| Gemini-2.0-Flash (Original) | 74.17% | 73.74% | 64.54% | 63.07% |
| Gemini-2.0-Flash (Full Permutation) | 75.67% | 75.56% | 67.97% | 65.52% |
Limitations
The study is bounded by filtering out audio samples longer than 180 seconds and restricting evaluations to four-option multiple-choice questions. Permutation-based mitigation strategies introduce significant computational overhead at inference time, multiplying API or compute costs by the number of permutations. Additionally, synthetic TTS audio generation may not fully capture the acoustic diversity of natural human speech corpora.
Why read this
Researchers and benchmark designers evaluating LALMs should read this paper to understand how severely option ordering distorts model rankings and to adopt permutation-based evaluation protocols for trustworthy results.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Standardized evaluation frameworks for audio-language models, speech-based automated testing systems, and robust multi-modal benchmark design.
Institutions
National Taiwan University
Funding / 經費: Ministry of Education, Taiwan Centers of Excellence in Artificial Intelligence, NTU Artificial Intelligence Center of Research Excellence
Related
- Robustness Assessment of Large Audio Language Models in Multiple-choice Evaluation — same problem · relatedness 2.7/3
- CoRE: Contrastive Evidence-Aware Rescoring for Multiple-Choice Audio Question Answering — same problem · relatedness 2.2/3
- ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge — same problem · relatedness 2.2/3
- A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models — same problem · relatedness 2.2/3
- MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models — same problem · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1025