TL;DR — This paper introduces a reasoning-guided ordinal speech emotion recognition framework that adapts large audio-language models (LALMs) for pairwise emotional comparisons, achieving superior preference accuracy while using only 5% of the training data required by conventional self-supervised baselines.
Key contributions
- Formulates speech emotion recognition as an explicit pairwise comparative reasoning task using LALMs rather than absolute regression or classification.
- Proposes a structured reasoning trace generation pipeline combining semantic audio descriptions with discretized GeMAPS acoustic low-level descriptors (LLDs).
- Applies supervised fine-tuning (SFT) and direct preference optimization (DPO) conditioned on reasoning traces (CoT) to improve both preference prediction and interpretability.
- Demonstrates strong data efficiency, outperforming conventional SSL ranking models using only 10k pairs (5% of baseline data volume).
Problem
Large audio-language models (LALMs) are predominantly designed for single-audio inference and struggle with multi-audio comparative reasoning across emotional, prosodic, or environmental dimensions. While psychological studies show that humans evaluate emotions more reliably via relative pairwise comparisons than absolute scores, existing emotion preference learning approaches learn implicitly from latent representations without modeling underlying acoustic cues. This leaves LALMs unable to provide transparent, interpretable justifications for why one speech utterance exhibits higher emotional intensity than another.
Method
The framework takes a pair of speech clips (xA, xB) alongside a task prompt specifying the emotional attribute (arousal, valence, or dominance). To construct reasoning traces, semantic captions are generated using Qwen3-Omni-Captioner, while 18 acoustic low-level descriptors (LLDs) plus their means and standard deviations (yielding a 36-dimensional GeMAPS representation) are extracted, normalized across the training set, and mapped to qualitative levels (low, medium, high) such as pitch, loudness, and vocal stability. A large reasoning model (Qwen3-Next-80B) synthesizes these into a concise reasoning trace (under 5 sentences) paired with the correct decision (y+), with automated regeneration conditioned on correct labels if the initial trace fails.
For alignment, the base model Qwen2.5-Omni-3B is adapted using LoRA (rank r = 64, scaling factor alpha = 64) applied to all linear layers. The system is trained via Supervised Fine-Tuning (SFT) to generate both the reasoning trace and final answer (r+, y+). Furthermore, Direct Preference Optimization (DPO) is utilized by pairing correct reasoning-answer tuples (r+, y+) against incorrect ones (r-, y-) generated by prompting the reasoning model with incorrect decisions, enforcing robust preference separation and suppressing hallucinations.
Experimental setup
Experiments use the MSP-Podcast v2.0 corpus (409 hours; 169k training, 34k dev, 46k test segments) for primary training and evaluation, constructing 10k training pairs and 3k test pairs per attribute using an absolute consensus score difference threshold > 1 on a 1–7 Likert scale. Cross-domain generalization is evaluated on the BIIC-Podcast corpus (Mandarin) and the WHiSER corpus (President Nixon Oval Office recordings, 1971–1973). Baselines include WavLM and HuBERT features combined with RankNet (trained on 240k pairs), and RankList. Models are evaluated using attribute-specific and average preference accuracy.
Results
On MSP-Podcast test pairs, the proposed DPO-CoT model achieves an average preference accuracy of 0.881, outperforming WavLM + RankNet (0.784), HuBERT + RankNet (0.765), and RankList (0.796), while using substantially fewer training pairs (10k vs 240k). Zero-shot Qwen2.5-Omni-3B performs poorly at 0.637 average accuracy, but standard SFT boosts this to 0.875 and DPO to 0.879. In cross-domain evaluations on WHiSER, DPO achieves 0.911 average preference accuracy compared to 0.764 for WavLM + RankNet. In cross-emotion generalization (training on arousal only and testing across dimensions), DPO-CoT achieves an average accuracy of 0.785, substantially mitigating performance drops on valence.
| Model | Arousal | Valence | Dominance | Avg |
|---|---|---|---|---|
| WavLM + RankNet | 0.792 | 0.806 | 0.753 | 0.784 |
| HuBERT + RankNet | 0.781 | 0.773 | 0.742 | 0.765 |
| RankList [57] | 0.808 | 0.813 | 0.767 | 0.796 |
| Qwen2.5-Omni-3B (Zero-shot) | 0.658 | 0.707 | 0.547 | 0.637 |
| + SFT | 0.881 | 0.878 | 0.867 | 0.875 |
| + DPO-CoT | 0.887 | 0.890 | 0.867 | 0.881 |
Limitations
The framework relies on pre-computed textual captions and GeMAPS heuristic features to anchor the reasoning trace, limiting end-to-end native acoustic reasoning. Evaluation is restricted to English, Mandarin, and historical archival audio datasets, leaving low-resource dialects and highly noisy real-world acoustic environments underexplored. The generation of reasoning traces introduces additional inference latency compared to direct classification heads.
Why read this
Researchers building multimodal audio-language models and speech emotion recognition systems should read this paper to learn how to inject acoustic-grounded reasoning traces and direct preference optimization into LALMs for robust multi-audio comparison.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Speech emotion recognition, psychiatric assessment tools, empathetic conversational agents, and multi-utterance affective analysis.
Institutions
Carnegie Mellon University, University of Texas at Dallas, NVIDIA
Related
- AcoustEmo: An Utterance-Aware Acoustic Q-Former for Open-Vocabulary Emotion Reasoning — same problem · relatedness 2.3/3
- Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction — same problem · relatedness 2.3/3
- Segment-wise Embedding based Graph Attention Network for Effective Speech Emotion Recognition — same problem · relatedness 2.2/3
- ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge — same problem · relatedness 2.2/3
- CHUCKLE - When Humans Teach AI to Learn Emotions the Easy Way — same problem · relatedness 2.1/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2935