---
id: lopez26_interspeech
title: Robustness Assessment of Large Audio Language Models in Multiple-choice
  Evaluation
authors:
  - Fernando López
  - Santosh Kesiraju
  - Jordi Luque
year: 2026
doi: 10.21437/Interspeech.2026-2503
isca_url: https://www.isca-archive.org/interspeech_2026/lopez26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/lopez26_interspeech.pdf
session: "Audio & Speech Language Models: Evaluation, Representations, and
  Emerging Capabilities"
topics:
  - speech-llm
  - evaluation
  - spoken-language-understanding
category: speech-llm-dialogue
institutions:
  - Telefonica
  - Universidad Autonoma de Madrid
  - Brno University of Technology
funding:
  - European Union's Horizon 2020
  - Ministry of Education, Youth and Sports of the Czech Republic
code:
  url: https://github.com/ferugit/mcqa-lalms-robustness
  stars: 1
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: lopez26_interspeech
  category: speech-llm-dialogue
  institutions:
    - Telefonica
    - Universidad Autonoma de Madrid
    - Brno University of Technology
  code: https://github.com/ferugit/mcqa-lalms-robustness
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-2503
  pdf: https://www.isca-archive.org/interspeech_2026/lopez26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/lopez26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/lopez26_interspeech/markdown.md
---

# Robustness Assessment of Large Audio Language Models in Multiple-choice Evaluation

*Fernando López, Santosh Kesiraju, Jordi Luque*

[PDF](https://www.isca-archive.org/interspeech_2026/lopez26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/lopez26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-2503)

**Category:** `speech-llm-dialogue`

**TL;DR** — A systematic robustness evaluation of large audio language models (LALMs) under multiple-choice question answering (MCQA) reveals that models are highly sensitive to linguistic variations, choice ordering, and distractor phrasing, while text-only controls expose significant language bias.

## Key contributions

- Quantified widespread language bias: text-only LLMs without audio input achieve up to 48.3% accuracy on MMAU (22.6 points above chance), proving existing benchmarks contain exploitable textual shortcuts.
- Demonstrated that minor prompt variations drastically shift performance, with distractor rephrasing causing accuracy standard deviations of up to 13.7%.
- Identified a length bias where LALMs disproportionately select longer answer candidates, peaking at nearly 78% selection rates when the longest option is correct.
- Proposed a mixed-perturbation evaluation protocol using the Correctness Rate (CoR) metric to efficiently capture model stability and robustness at practical compute costs.

## Problem

Current large audio language models are predominantly evaluated using multiple-choice question answering frameworks that report a single aggregate accuracy score. However, this methodology fails to verify whether models genuinely reason from the audio signal or exploit superficial linguistic hints, prompt framing, and formatting artifacts. Prior MCQA setups in both text LLMs and LALMs ignore sensitivity to choice permutations and paraphrasing, leading to inflated performance metrics that obscure true model capabilities and limitations.

## Method

The study evaluates four open-source LALMs—Audio Flamingo 2 (3.2B cross-attention), Audio Flamingo 3 (8.4B self-attention), Qwen2.5-Omni-7B (7B self-attention), and Kimi-Audio-7B-Instruct (7B self-attention)—under a fixed-audio perturbation protocol. The evaluation isolates linguistic sensitivity through five perturbations: choice ordering (all 24 permutations), question rephrasing (7 variants generated by gemini-2.5-flash and gemma-3-12b-it), ground-truth answer rephrasing (7 variants), distractor rephrasing (7 variants), and a mixed-perturbation setting where modifications are applied independently with a probability of 0.5.

Models are queried using a standardized template with greedy decoding, and their outputs are parsed to extract selected option letters and text. To evaluate performance beyond mean accuracy, the authors employ the Consistency Rate (CR) to measure internal response agreement across invariant audio inputs, and the Correctness Rate (CoR), which treats a question as correct only if answered accurately across all perturbed versions. Text-only controls (Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and gemma-3-27B-it) receive identical text prompts without audio to quantify language priors.

## Experimental setup

Evaluations span three established benchmarks: MMAU-v05.15.25 test-mini subset (1k samples across speech, sound, and music), MMAR (1k samples across speech, sound, music, and mixed sources), and MMSU (5k speech samples focused on reasoning). Baselines include random choice, text-only LLMs, and legacy audio models (LTU-AS, SALMONN, Qwen-Audio-Chat). Performance is measured using mean accuracy, standard deviation, min/max ranges, consistency rate, and correctness rate.

## Results

On the default setup, Audio Flamingo 3 leads with 73.3% average accuracy on MMAU, followed closely by Qwen2.5-Omni-7B (68.2%) and Kimi-Audio-7B-Instruct (71.1%). However, performance degrades sharply under linguistic perturbations: distractor rephrasing causes massive drops, with Audio Flamingo 2's accuracy plummeting from 62.7% down to an average of 42.7% (std 13.7%). While Audio Flamingo 3 achieves top absolute accuracy, Qwen2.5-Omni-7B demonstrates superior robustness against distractor wording, securing the highest Correctness Rate (CoR) across benchmarks under distractor perturbations.

Ablations on option length reveal a strong shortcut reliance: although only 45.05% of dataset choices are the longest option, Audio Flamingo 3 and Kimi-Audio select the longest choice over 50% of the time generally, and up to 77.94% and 76.41% of the time respectively when the longest option happens to be the ground truth.

| System / Condition | MMAU Accuracy (%) | MMAR Accuracy (%) | MMSU Accuracy (%) |
|---|---|---|---|
| Random Choice | 25.7 | 33.0 | 25.4 |
| gemma-3-27B-it (Text-only) | 48.3 | 35.8 | 38.9 |
| Audio Flamingo 2 (Default) | 62.7 | 45.3 | 41.8 |
| Audio Flamingo 3 (Default) | 73.3 | 58.5 | 61.0 |
| Qwen2.5-Omni-7B (Default) | 68.2 | 59.0 | 62.4 |
| Kimi-Audio-7B-Instruct (Default) | 71.1 | 54.0 | 58.4 |

## Limitations

The study isolates linguistic robustness while keeping the audio signal constant, meaning signal-level acoustic perturbations (such as noise or duration shifts) are outside its current scope. The evaluation is limited to four open-source models and three benchmarks, and the generated paraphrases, while manually validated on a subset of 620 samples, may not capture all real-world linguistic nuances or cross-lingual variations.

## Why read this

Speech and ML researchers designing evaluation benchmarks or training multimodal audio language models should read this to understand how easily MCQA results can be manipulated by text formatting, option ordering, and language bias. It provides actionable evaluation protocols like the Correctness Rate and mixed-perturbation testing to measure true model robustness.

## Code

- https://github.com/ferugit/mcqa-lalms-robustness

## Applications

Improving the reliability and robustness of automated audio evaluation benchmarks, speech-based question-answering systems, and multi-modal conversational assistants.

## Institutions / 機構

Telefonica, Universidad Autonoma de Madrid, Brno University of Technology

**Funding / 經費:** European Union's Horizon 2020, Ministry of Education, Youth and Sports of the Czech Republic

## Related

- [Hearing the Order: Investigating Position Bias in Large Audio-Language Models](lin26c_interspeech.md) — same problem · relatedness 2.7/3
- [CoRE: Contrastive Evidence-Aware Rescoring for Multiple-Choice Audio Question Answering](zhang26f_interspeech.md) — same problem · relatedness 2.3/3
- [All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation](foo26_interspeech.md) — same problem · relatedness 2.2/3
- [Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models](seth26_interspeech.md) — same problem · relatedness 2.2/3
- [UG-Bench: A Comprehensive Benchmark for Evaluating Large Audio-Language Models](zhou26c_interspeech.md) — same problem · relatedness 2.1/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
