All papers
Speech LLMs & dialogueFull-paper digest

Robustness Assessment of Large Audio Language Models in Multiple-choice Evaluation

Fernando López, Santosh Kesiraju, Jordi Luque

Code & resourcesgithub.com/ferugit/mcqa-lalms-robustness

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.0 KB · Ready to paste

Preview copied content

TL;DR — A systematic robustness evaluation of large audio language models (LALMs) under multiple-choice question answering (MCQA) reveals that models are highly sensitive to linguistic variations, choice ordering, and distractor phrasing, while text-only controls expose significant language bias.

Key contributions

  • Quantified widespread language bias: text-only LLMs without audio input achieve up to 48.3% accuracy on MMAU (22.6 points above chance), proving existing benchmarks contain exploitable textual shortcuts.
  • Demonstrated that minor prompt variations drastically shift performance, with distractor rephrasing causing accuracy standard deviations of up to 13.7%.
  • Identified a length bias where LALMs disproportionately select longer answer candidates, peaking at nearly 78% selection rates when the longest option is correct.
  • Proposed a mixed-perturbation evaluation protocol using the Correctness Rate (CoR) metric to efficiently capture model stability and robustness at practical compute costs.

Problem

Current large audio language models are predominantly evaluated using multiple-choice question answering frameworks that report a single aggregate accuracy score. However, this methodology fails to verify whether models genuinely reason from the audio signal or exploit superficial linguistic hints, prompt framing, and formatting artifacts. Prior MCQA setups in both text LLMs and LALMs ignore sensitivity to choice permutations and paraphrasing, leading to inflated performance metrics that obscure true model capabilities and limitations.

Method

The study evaluates four open-source LALMs—Audio Flamingo 2 (3.2B cross-attention), Audio Flamingo 3 (8.4B self-attention), Qwen2.5-Omni-7B (7B self-attention), and Kimi-Audio-7B-Instruct (7B self-attention)—under a fixed-audio perturbation protocol. The evaluation isolates linguistic sensitivity through five perturbations: choice ordering (all 24 permutations), question rephrasing (7 variants generated by gemini-2.5-flash and gemma-3-12b-it), ground-truth answer rephrasing (7 variants), distractor rephrasing (7 variants), and a mixed-perturbation setting where modifications are applied independently with a probability of 0.5.

Models are queried using a standardized template with greedy decoding, and their outputs are parsed to extract selected option letters and text. To evaluate performance beyond mean accuracy, the authors employ the Consistency Rate (CR) to measure internal response agreement across invariant audio inputs, and the Correctness Rate (CoR), which treats a question as correct only if answered accurately across all perturbed versions. Text-only controls (Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and gemma-3-27B-it) receive identical text prompts without audio to quantify language priors.

Experimental setup

Evaluations span three established benchmarks: MMAU-v05.15.25 test-mini subset (1k samples across speech, sound, and music), MMAR (1k samples across speech, sound, music, and mixed sources), and MMSU (5k speech samples focused on reasoning). Baselines include random choice, text-only LLMs, and legacy audio models (LTU-AS, SALMONN, Qwen-Audio-Chat). Performance is measured using mean accuracy, standard deviation, min/max ranges, consistency rate, and correctness rate.

Results

On the default setup, Audio Flamingo 3 leads with 73.3% average accuracy on MMAU, followed closely by Qwen2.5-Omni-7B (68.2%) and Kimi-Audio-7B-Instruct (71.1%). However, performance degrades sharply under linguistic perturbations: distractor rephrasing causes massive drops, with Audio Flamingo 2's accuracy plummeting from 62.7% down to an average of 42.7% (std 13.7%). While Audio Flamingo 3 achieves top absolute accuracy, Qwen2.5-Omni-7B demonstrates superior robustness against distractor wording, securing the highest Correctness Rate (CoR) across benchmarks under distractor perturbations.

Ablations on option length reveal a strong shortcut reliance: although only 45.05% of dataset choices are the longest option, Audio Flamingo 3 and Kimi-Audio select the longest choice over 50% of the time generally, and up to 77.94% and 76.41% of the time respectively when the longest option happens to be the ground truth.

System / ConditionMMAU Accuracy (%)MMAR Accuracy (%)MMSU Accuracy (%)
Random Choice25.733.025.4
gemma-3-27B-it (Text-only)48.335.838.9
Audio Flamingo 2 (Default)62.745.341.8
Audio Flamingo 3 (Default)73.358.561.0
Qwen2.5-Omni-7B (Default)68.259.062.4
Kimi-Audio-7B-Instruct (Default)71.154.058.4

Limitations

The study isolates linguistic robustness while keeping the audio signal constant, meaning signal-level acoustic perturbations (such as noise or duration shifts) are outside its current scope. The evaluation is limited to four open-source models and three benchmarks, and the generated paraphrases, while manually validated on a subset of 620 samples, may not capture all real-world linguistic nuances or cross-lingual variations.

Why read this

Speech and ML researchers designing evaluation benchmarks or training multimodal audio language models should read this to understand how easily MCQA results can be manipulated by text formatting, option ordering, and language bias. It provides actionable evaluation protocols like the Correctness Rate and mixed-perturbation testing to measure true model robustness.

Code

Applications

Improving the reliability and robustness of automated audio evaluation benchmarks, speech-based question-answering systems, and multi-modal conversational assistants.

Institutions

Telefonica, Universidad Autonoma de Madrid, Brno University of Technology

Funding / 經費: European Union's Horizon 2020, Ministry of Education, Youth and Sports of the Czech Republic

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2503