All papers
Speech LLMs & dialogueFull-paper digest

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

Alexander Polok, Samuele Cornell, Sathvik Udupa, Honza Černocký, Shinji Watanabe, Lukáš Burget

Code & resourcesgithub.com/BUTSpeechFIT/Dixtral

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.8 KB · Ready to paste

Preview copied content

TL;DR — Dixtral extends Spoken Large Language Models (SLMs) to far-field multi-talker audio by conditioning a Whisper-based acoustic encoder on diarization masks while keeping the LLM decoder frozen. It outperforms Gemini 3.0 Flash, VibeVoice, and Voxtral Mini Transcribe V2 on speaker-attributed transcription by 29.0%, 19.8%, and 16.0% absolute cpWER respectively.

Key contributions

  • Introduces Diarization-Conditioned Spoken LLMs, bypassing Serialized Output Training (SOT) to avoid catastrophic forgetting and vocabulary expansion.
  • Proposes Dixtral, combining a Diarization-Conditioned Whisper (DiCoW) encoder with a frozen Voxtral/Ministral 3B decoder for target-speaker extraction, transcription, QA, and summarization.
  • Releases NSF-QA, a long-form multi-speaker question-answering and summarization benchmark built on NOTSOFAR-1 covering content, emotion, and speaker gender.
  • Demonstrates that independent target-speaker extraction scales generation with O(S * N^2) complexity rather than O((SN)^2) sequence-length penalties in SOT models.

Problem

Spoken LLMs typically process multi-talker audio using Serialized Output Training (SOT), which serializes transcripts by onset time, expands the LLM vocabulary with special tokens, and requires heavy decoder fine-tuning. This causes severe distributional mismatch and catastrophic forgetting of core reasoning and QA capabilities. Modular pipelines solve transcription well but lack zero-shot generalization to cross-modal tasks like summarization and conversational QA.

Method

Dixtral replaces the acoustic encoder of an SLM (Voxtral Mini 3B, using a Whisper large-v3 encoder and a Ministral 3B decoder) with a diarization-conditioned encoder derived from DiCoW. Frame-level speaker activity probabilities (Silence, Target, Non-target, Overlap - STNO) modulate the internal representations of each encoder layer via learnable diagonal affine transformation matrices using Frame-Level Diarization-Dependent Transformations (FDDT). The conditioned acoustic representations are passed through a two-layer MLP modality adapter with GELU activations to project them into the LLM embedding space. Text prompts and projected audio embeddings are concatenated as a prefix sequence, feeding a frozen Ministral 3B decoder that generates the target response autoregressively using standard cross-entropy loss over the generated tokens.

During training, both the LLM decoder and modality adapter are completely frozen; only the acoustic encoder layers and FDDT modules are updated. For multi-task scaling, the model is trained on 8x 24GB A5000 GPUs using bfloat16 precision, gradient checkpointing, and gradient accumulation of 4 steps (global batch size 32) for 20,000 steps with a peak learning rate of 6e-5 (5,000 warmup steps, cosine decay). Long-form QA fine-tuning is performed on H100 GPUs.

Experimental setup

Evaluated on four multi-talker transcription datasets: NOTSOFAR-1 (NSF-1), AMI, LibriSpeechMix (LSMix), and Mixer6 (MX-6, out-of-domain evaluation). Transcription is measured via Concatenated Minimum-Permutation Word Error Rate (cpWER). QA and summarization are evaluated on the novel NSF-QA benchmark using Gemini 2.5 Flash as an LLM judge for accuracy and ROUGE-L for summarization. Compared against DiCoW v3.3, VibeVoice, Voxtral MTv2, and Gemini 3.0 Flash.

Results

On speaker-attributed transcription across all datasets, Dixtral achieves a macro-average cpWER of 15.4%, outperforming Gemini 3.0 Flash (44.4%), VibeVoice (35.2%), and Voxtral Mini Transcribe V2 (31.4%). Specifically on NSF-1, Dixtral scores 29.1% cpWER compared to Gemini's 39.1% and VibeVoice's 35.8%. On the NSF-QA benchmark, zero-shot Dixtral achieves 54.6% Content QA accuracy (matching far-field Gemini at 55.1%) and 24.4 ROUGE-L on summarization. When fine-tuned on NSF-QA, Dixtral reaches 73.0% Content QA, 47.6% Emotion QA, 95.5% Gender QA, and 41.4 ROUGE-L, surpassing both close-talk Voxtral and Gemini.

SystemNSF-1AMI SmallLSMix 1LSMix 2LSMix 3MX-6 CH4Average
Voxtral MTv254.442.32.028.242.319.431.4
VibeVoice35.833.72.150.872.816.035.2
DiCoW v3.326.618.61.83.121.711.914.0
Gemini 3.0 Flash39.156.34.523.384.758.344.4
Dixtral29.119.82.13.623.514.415.4

Limitations

The approach relies entirely on external diarization systems (such as DiariZen), meaning errors in upstream diarization directly cascade into target-speaker extraction failures. Joint end-to-end training with diarization has not yet been explored, and task-specific fine-tuning for QA/summarization can cause a trade-off that slightly degrades verbatim transcription accuracy on out-of-domain data like Mixer-6.

Why read this

Speech and ML researchers working on far-field multi-talker audio or Spoken LLMs should read this paper to see how target-speaker extraction via encoder conditioning avoids SOT-related catastrophic forgetting while outperforming proprietary models like Gemini on speaker-attributed speech tasks.

Code

Applications

Automated meeting transcription and minutes generation, multi-speaker conversational AI assistants, far-field smart home voice interfaces, and multi-talker audio question-answering systems.

Institutions

Brno University of Technology, Carnegie Mellon University

Funding / 經費: Ministry of Education, Youth and Sports of the Czech Republic, Brno Ph.D. Talent Scholarship Programme, e-INFRA CZ

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-445