All papers
Paralinguistics & emotionFull-paper digest

ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding

Nishit Anand, Jiaqi Su, Ke Chen, Yunyun Wang, Dinesh Manocha, Ramani Duraiswami, Rithesh Kumar, Zeyu Jin

Code & resourcesnishitanand.github.io/paralinguistic-understanding-llm

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

11.9 KB · Ready to paste

Preview copied content

TL;DR — ParA-LLM is an Audio LLM trained with a two-stage curriculum on 1.2M simulated and annotated QA pairs to master 22 paralinguistic and acoustic characteristics, outperforming GPT-4o-Audio by 7.5% overall on the newly introduced ParA-Bench.

Key contributions

  • Defined a structured taxonomy of 22 objective and subjective paralinguistic characteristics spanning 10 acoustic properties, 7 speaker-intrinsic traits, and 5 utterance-level speech attributes.
  • Curated a large-scale dataset of over 1.2M Audio-QA pairs via an acoustic simulation engine (mixing RIRs and environmental noise) and LLM-assisted generation.
  • Developed ParA-LLM via a two-stage curriculum (atomic single-attribute followed by multi-attribute joint reasoning).
  • Introduced ParA-Bench, a 6,000 multiple-choice question benchmark covering speaker-speech, acoustic, and mixed evaluation categories.

Problem

While contemporary Audio LLMs (such as Qwen2-Audio, Voxtral, and GPT-4o-Audio) excel at verbal semantic tasks like transcription, they largely fail to understand paralinguistic dimensions such as acoustic conditions, speaker traits, and expressive variations. Humans achieve 78% accuracy on paralinguistic benchmarks like ParA-Bench, whereas frontier models score only around 36%. Specialized models only classify isolated attributes and lack free-form compositional reasoning over multiple interacting characteristics simultaneously.

Method

ParA-LLM is initialized from Qwen2-Audio-7B-Instruct and trained using LoRA (rank 128, alpha 256, dropout 0.1) applied to the audio encoder, multimodal projector, and LLM backbone. The training pipeline uses a two-stage curriculum: Stage 1 trains on 688K atomic single-attribute QA pairs (built from 306K unique samples) for 1 epoch at a learning rate of 5e-5, while Stage 2 continues on 513K multi-attribute QA pairs (from 217K samples) for 1 epoch at 4e-5.

The training data is synthesized by augmenting clean speech from EARS, Emilia, Expresso, and VoxCeleb with 15 reverb types (via MIT IR Survey and EchoThief RIRs), 21 noise types (via TAU and urban sound datasets), and post-production effects (clipping, compression, overdrive). Characteristics are mapped along objective signal metrics (e.g., DRR, RT60, SNR, STOI) and consistent natural-language labels for speaker timbre, accent, nasality, and emotion. Inference supports free-form, multi-attribute question answering over audio inputs.

Experimental setup

Evaluated on ParA-Bench (6,000 multiple-choice questions across speaker-speech, acoustic, and mixed domains), MMAU-Pro Speech, and MMAR Speech. Compared against baselines including Qwen2-Audio, Voxtral-24B, Audio Flamingo 3, Mellow, R1-AQA, Qwen2.5-Omni (3B/7B), and GPT-4o-Audio. Training utilized 8 A100 GPUs with a global batch size of 128 using the AdamW optimizer and cosine learning rate scheduling.

Results

ParA-LLM achieves an overall ParA-Bench accuracy of 43.53%, surpassing Voxtral (38.80%) by 4.73% and GPT-4o-Audio (36.03%) by 7.50%. It scores highest in speaker-speech (55.85%) and mixed categories (39.95%), though GPT-4o-Audio retains the lead in the purely acoustic category (41.85% vs ParA-LLM's 34.80%). On broader evaluation sets, the two-stage curriculum boosts MMAR Speech performance from 35.37% (baseline) to 42.86% (Stage 2).

ModelSpeaker-SpeechAcousticMixedOverall
Qwen2 Audio34.8023.4525.9028.05
Voxtral51.9529.2035.2538.80
GPT-4o-Audio32.0041.8534.2536.03
ParA-LLM (Ours)55.8534.8039.9543.53

Limitations

The evaluation relies on synthetic simulation mixtures for acoustic conditions which may not fully capture every real-world acoustic anomaly. The benchmark and dataset distribution are currently English-centric or drawn from standard corpora, and model performance drops on pure acoustic-only attribute queries compared to proprietary frontier systems like GPT-4o-Audio.

Why read this

Speech and ML engineers working on expressive speech synthesis, acoustic editing, or fine-grained audio reasoning should read this to understand how curriculum learning and structured taxonomies can systematically bridge the paralinguistic capability gap in Audio LLMs.

Code

Applications

Scalable automatic metadata annotation for speech corpora, fine-grained attribute control for text-to-speech (TTS) systems, and acoustic space captioning for text-conditioned room impulse response generation (Text2IR).

Institutions

Adobe Research, University of Maryland, College Park, OpenAI

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3015