All papers
Resources & evaluationFull-paper digest

MSU-Bench: Towards Understanding the Conversational Multi-Speaker Scenarios

Zhaokai Sun, Shuai Wang, Zhennan Lin, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.4 KB · Ready to paste

Preview copied content

TL;DR — MSU-Bench is a diagnostic benchmark designed to evaluate large audio language models on speaker-centric understanding across realistic multi-speaker conversations, covering 16 tasks and 2,300 verified QA instances. Evaluation of nine models reveals a performance range from 0.19 to 0.77 exact-match accuracy, highlighting major bottlenecks in temporal grounding and speaker attribution.

Key contributions

  • Introduces MSU-Bench, a two-tier diagnostic benchmark containing 16 speaker-centric tasks across five capabilities and 2,300 QA instances.
  • Combines eight diverse conversational and media-style corpora (totaling over 640 hours) to test robustness in meetings, telephone calls, podcasts, and movies in both Chinese and English.
  • Proposes five distinct speaker-referencing schemes (No Index, Time Index, Transcript Index, Speaker Index, Complex Index) to systematically probe speaker grounding abilities.
  • Implements a Gemini-assisted annotation and QA pipeline with human-in-the-loop verification, achieving 96%-98% human-ground truth agreement.

Problem

Current speech benchmarks primarily evaluate single-speaker scenarios or isolate individual subtasks like standard ASR, simple emotion recognition, or standard speaker verification. Real-world audio interactions feature rapid turn-taking, overlaps, and background noise, requiring models to simultaneously track speaker identity, consistency, dialogue structure, and cross-turn reasoning. Prior evaluation infrastructure lacks the capacity to diagnose these multi-speaker conversational dynamics, leaving the precise failure modes of large audio language models (LALMs) unclear.

Method

MSU-Bench establishes a hierarchical task structure divided into Tier 1 (Speaker Grounding and Identification, encompassing tasks like Speaker Retrieval, Speaker Counting, and Accent/Age/Gender/Emotion Recognition) and Tier 2 (Multi-Speaker Dialogue Reasoning, encompassing Background Inference, Role Identification, Dialogue Act Recognition, and Emotion Interaction Reasoning). Data is sourced from 8 corpora including Chinese/English telephone dialogues, AliMeeting (12h), CHiME-6 (12h), podcasts (97h), and movies (600h), preprocessed into clips ranging from 1 to 5 minutes.

The benchmark construction pipeline utilizes Gemini for dialogue quality assessment, Volcano API for initial diarization and transcripts, and Gemini for paralinguistic and identity metadata annotation. Gemini generates four-option multiple-choice QA pairs based on these annotations, using three diagnostic distractors: wrong-speaker (WS), hallucination (HAL), and unknown (UNK). Trained human annotators subsequently review, revise, or discard ambiguous samples to guarantee validity. Inference is evaluated zero-shot across nine speech-language models using identical instruction templates that demand a single option letter (A/B/C/D), measured via exact-match accuracy.

Experimental setup

Evaluated nine speech-language models: six open-source systems (Qwen2.5-Omni, Qwen3-Omni, AudioFlamingo-3, Kimi-Audio, StepAudio2, MiMoAudio) and three closed-source Gemini systems (Gemini-2.5-Flash, Gemini-2.5-Pro, Gemini-3-Flash). Evaluated on 2,300 four-option QA instances derived from 8 datasets spanning telephone conversations, meetings, podcasts, and movies in Chinese and English. Metrics include exact-match accuracy across tiers, tasks, and referencing schemes, along with diagnostic error-type composition.

Results

Exact-match accuracy spans widely from 0.19 (Qwen2.5-Omni) to 0.77 (Gemini-3-Flash). Closed-source models strictly dominate open-source counterparts: Gemini-3-Flash leads overall with 0.77 accuracy (0.73 Tier 1, 0.84 Tier 2), followed by Gemini-2.5-Pro (0.70) and Gemini-2.5-Flash (0.69). Among open-source models, MiMoAudio performs best with 0.56 overall (0.52 Tier 1, 0.64 Tier 2), outperforming StepAudio2 (0.44), Kimi-Audio (0.43), AudioFlamingo-3 (0.39), and Qwen3-Omni (0.39).

Ablations over speaker-referencing schemes show that Time Index poses the hardest challenge (e.g., Gemini-3-Flash drops to 0.64 on Tier 1 and 0.76 on Tier 2), whereas Complex Index boosts performance (0.84 Tier 1, 0.92 Tier 2 for Gemini-3-Flash). Error analysis reveals that weaker open-source models frequently default to the unknown (UNK) option under high demands (e.g., Qwen3-Omni exhibits a 0.40 UNK rate on Tier 2), whereas advanced models commit errors predominantly via wrong-speaker (WS) attribution (e.g., Gemini-3-Flash reaches a 0.67 WS error rate on Tier 2).

SystemsTier 1 AccTier 2 AccOverall Avg
Qwen2.5-Omni0.190.210.19
AudioFlamingo-30.400.380.39
Qwen3-Omni0.400.380.39
Kimi-Audio0.410.470.43
StepAudio20.440.460.44
MiMoAudio0.520.640.56
Gemini-2.5-Flash0.640.780.69
Gemini-2.5-Pro0.670.820.70
Gemini-3-Flash0.730.840.77

Limitations

The evaluation is constrained to 2,300 English and Mandarin Chinese QA instances, leaving low-resource languages untested. Evaluating via multiple-choice tasks, while robust and deterministic, does not fully capture open-ended conversational generation behaviors or streaming dialogue scenarios. Additionally, while human-in-the-loop verification was enforced, utilizing Gemini for initial QA synthesis may introduce subtle dataset biases.

Why read this

Researchers and engineers building conversational large audio language models should read this paper to understand the specific failure boundaries of current models in multi-speaker grounding and interaction reasoning. It provides an off-the-shelf diagnostic testbed to benchmark architectural improvements in temporal speaker tracking.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Diagnostic evaluation of conversational speech foundation models, multi-speaker meeting assistant refinement, and robust audio-language model architectural development.

Institutions

Northwestern Polytechnical University, Nanjing University, Shenzhen Loop Area Institute, Li Auto

Funding / 經費: National Natural Science Foundation of China

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3344