---
id: naini26_interspeech
title: "Comparative Reasoning: Making an Audio Language Model Better at
  Comparing Emotions"
authors:
  - Abinay Reddy Naini
  - Jaeyeon Kim
  - Chao-Han Huck Yang
  - Shinji Watanabe
  - Carlos Busso
year: 2026
doi: 10.21437/Interspeech.2026-2935
isca_url: https://www.isca-archive.org/interspeech_2026/naini26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/naini26_interspeech.pdf
session: Multimodal Emotion Recognition
topics:
  - speech-emotion-recognition
  - speech-llm
  - self-supervised
category: paralinguistics-emotion
labels:
  - low-resource
  - self-supervised
institutions:
  - Carnegie Mellon University
  - University of Texas at Dallas
  - NVIDIA
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: naini26_interspeech
  category: paralinguistics-emotion
  labels:
    - low-resource
    - self-supervised
  institutions:
    - Carnegie Mellon University
    - University of Texas at Dallas
    - NVIDIA
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-2935
  pdf: https://www.isca-archive.org/interspeech_2026/naini26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/naini26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/naini26_interspeech/markdown.md
---

# Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions

*Abinay Reddy Naini, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Carlos Busso*

[PDF](https://www.isca-archive.org/interspeech_2026/naini26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/naini26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-2935)

**Category:** `paralinguistics-emotion` · **Labels:** `low-resource`, `self-supervised`

**TL;DR** — This paper introduces a reasoning-guided ordinal speech emotion recognition framework that adapts large audio-language models (LALMs) for pairwise emotional comparisons, achieving superior preference accuracy while using only 5% of the training data required by conventional self-supervised baselines.

## Key contributions

- Formulates speech emotion recognition as an explicit pairwise comparative reasoning task using LALMs rather than absolute regression or classification.
- Proposes a structured reasoning trace generation pipeline combining semantic audio descriptions with discretized GeMAPS acoustic low-level descriptors (LLDs).
- Applies supervised fine-tuning (SFT) and direct preference optimization (DPO) conditioned on reasoning traces (CoT) to improve both preference prediction and interpretability.
- Demonstrates strong data efficiency, outperforming conventional SSL ranking models using only 10k pairs (5% of baseline data volume).

## Problem

Large audio-language models (LALMs) are predominantly designed for single-audio inference and struggle with multi-audio comparative reasoning across emotional, prosodic, or environmental dimensions. While psychological studies show that humans evaluate emotions more reliably via relative pairwise comparisons than absolute scores, existing emotion preference learning approaches learn implicitly from latent representations without modeling underlying acoustic cues. This leaves LALMs unable to provide transparent, interpretable justifications for why one speech utterance exhibits higher emotional intensity than another.

## Method

The framework takes a pair of speech clips (xA, xB) alongside a task prompt specifying the emotional attribute (arousal, valence, or dominance). To construct reasoning traces, semantic captions are generated using Qwen3-Omni-Captioner, while 18 acoustic low-level descriptors (LLDs) plus their means and standard deviations (yielding a 36-dimensional GeMAPS representation) are extracted, normalized across the training set, and mapped to qualitative levels (low, medium, high) such as pitch, loudness, and vocal stability. A large reasoning model (Qwen3-Next-80B) synthesizes these into a concise reasoning trace (under 5 sentences) paired with the correct decision (y+), with automated regeneration conditioned on correct labels if the initial trace fails.

For alignment, the base model Qwen2.5-Omni-3B is adapted using LoRA (rank r = 64, scaling factor alpha = 64) applied to all linear layers. The system is trained via Supervised Fine-Tuning (SFT) to generate both the reasoning trace and final answer (r+, y+). Furthermore, Direct Preference Optimization (DPO) is utilized by pairing correct reasoning-answer tuples (r+, y+) against incorrect ones (r-, y-) generated by prompting the reasoning model with incorrect decisions, enforcing robust preference separation and suppressing hallucinations.

## Experimental setup

Experiments use the MSP-Podcast v2.0 corpus (409 hours; 169k training, 34k dev, 46k test segments) for primary training and evaluation, constructing 10k training pairs and 3k test pairs per attribute using an absolute consensus score difference threshold > 1 on a 1–7 Likert scale. Cross-domain generalization is evaluated on the BIIC-Podcast corpus (Mandarin) and the WHiSER corpus (President Nixon Oval Office recordings, 1971–1973). Baselines include WavLM and HuBERT features combined with RankNet (trained on 240k pairs), and RankList. Models are evaluated using attribute-specific and average preference accuracy.

## Results

On MSP-Podcast test pairs, the proposed DPO-CoT model achieves an average preference accuracy of 0.881, outperforming WavLM + RankNet (0.784), HuBERT + RankNet (0.765), and RankList (0.796), while using substantially fewer training pairs (10k vs 240k). Zero-shot Qwen2.5-Omni-3B performs poorly at 0.637 average accuracy, but standard SFT boosts this to 0.875 and DPO to 0.879. In cross-domain evaluations on WHiSER, DPO achieves 0.911 average preference accuracy compared to 0.764 for WavLM + RankNet. In cross-emotion generalization (training on arousal only and testing across dimensions), DPO-CoT achieves an average accuracy of 0.785, substantially mitigating performance drops on valence.

| Model | Arousal | Valence | Dominance | Avg |
|---|---|---|---|---|
| WavLM + RankNet | 0.792 | 0.806 | 0.753 | 0.784 |
| HuBERT + RankNet | 0.781 | 0.773 | 0.742 | 0.765 |
| RankList [57] | 0.808 | 0.813 | 0.767 | 0.796 |
| Qwen2.5-Omni-3B (Zero-shot) | 0.658 | 0.707 | 0.547 | 0.637 |
| + SFT | 0.881 | 0.878 | 0.867 | 0.875 |
| + DPO-CoT | 0.887 | 0.890 | 0.867 | 0.881 |

## Limitations

The framework relies on pre-computed textual captions and GeMAPS heuristic features to anchor the reasoning trace, limiting end-to-end native acoustic reasoning. Evaluation is restricted to English, Mandarin, and historical archival audio datasets, leaving low-resource dialects and highly noisy real-world acoustic environments underexplored. The generation of reasoning traces introduces additional inference latency compared to direct classification heads.

## Why read this

Researchers building multimodal audio-language models and speech emotion recognition systems should read this paper to learn how to inject acoustic-grounded reasoning traces and direct preference optimization into LALMs for robust multi-audio comparison.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Speech emotion recognition, psychiatric assessment tools, empathetic conversational agents, and multi-utterance affective analysis.

## Institutions / 機構

Carnegie Mellon University, University of Texas at Dallas, NVIDIA

## Related

- [AcoustEmo: An Utterance-Aware Acoustic Q-Former for Open-Vocabulary Emotion Reasoning](zhang26ea_interspeech.md) — same problem · relatedness 2.3/3
- [Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction](yu26f_interspeech.md) — same problem · relatedness 2.3/3
- [Segment-wise Embedding based Graph Attention Network for Effective Speech Emotion Recognition](song26c_interspeech.md) — same problem · relatedness 2.2/3
- [ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge](jeon26d_interspeech.md) — same problem · relatedness 2.2/3
- [CHUCKLE - When Humans Teach AI to Learn Emotions the Easy Way](singh26b_interspeech.md) — same problem · relatedness 2.1/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
