All papers
Resources & evaluationFull-paper digest

Calibration-Reasoning Framework for Descriptive Speech Quality Assessment

Elizaveta Kostenok, Mathieu Salzmann, Milos Cernak

Code & resourcesgithub.com/KostenokLisa/calibration-reasoning-framework

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — A two-stage post-training framework combining a calibration phase with dimension-specific reinforcement learning (GRPO) enables Audio Large Language Models to accurately assess, describe, and temporally localize speech quality issues, achieving a state-of-the-art mean PCC of 0.71 and a 13% improvement in MOS prediction.

Key contributions

  • A two-stage training curriculum (Calibration followed by Reasoning) that prevents the dimension-score degradation typically caused by standard long-form generation fine-tuning.
  • Unfreezing and end-to-end training of the audio encoder during calibration, demonstrating a massive 0.12 average PCC boost compared to keeping encoders frozen.
  • A dimension-wise reward mechanism for Group Relative Policy Optimization (GRPO) using either an LLM-judge or parsed accuracy plus semantic similarity.
  • Superior temporal localization (Intersection over Union) and artifact classification for noise, distortion, and unnatural pauses compared to prior SFT and unified-reward RL methods.

Problem

Traditional speech quality assessment models predict scalar Mean Opinion Scores (MOS) as black boxes without explainability, while prior explainable audio LLM systems prioritize conversational fluency over diagnostic precision. Because quality assessment is missing from base pre-training mixtures, these models produce hallucinated dimension predictions, ungrounded reasoning, and degraded MOS accuracy. Existing multi-stage pipelines (like QualiSpeech-FT) suffer from catastrophic forgetting of numerical dimension scores during the second text-generation stage due to insufficient reasoning capacity or poorly structured reward functions.

Method

The framework builds on Audio Flamingo 3 (7-8B parameter range) and executes in two distinct stages. The first stage is Calibration via Supervised Fine-Tuning (SFT), where the audio encoder is unfrozen (trained end-to-end) and the model is optimized using Cross-Entropy loss to predict explicit perceptual scales [1, 5] across seven dimensions (naturalness, noise, distortion, listening effort, continuity, speed, and MOS).

The second stage is Reasoning via Group Relative Policy Optimization (GRPO). For a given audio input, the policy model samples a group of G=4 candidate responses. GRPO maximizes the probability of outputs yielding higher rewards while regularizing updates via Kullback-Leibler (KL) divergence against a frozen reference model to prevent reward hacking. Two alternative dimension-specific reward strategies are explored: an LLM-judge (using Qwen3) that evaluates dimension-wise generations, and an Accuracy + Semantic Similarity reward where numerical scores are evaluated exactly and text descriptions are scored using sentence transformer embeddings (all-MiniLM-L6-v2) mapped to [0,1].

Training uses LoRA with rank 64. Calibration uses a batch size of 8 for 10k iterations at a 1.5e-5 learning rate, SFT reasoning uses a batch size of 8 for 10k iterations at 1.5e-5, and GRPO uses a group size of 4, batch size of 4, for 200 iterations at a 5e-6 learning rate on two NVIDIA L40S GPUs.

Experimental setup

Evaluated on the QualiSpeech corpus consisting of 12,450 speech recordings (85% train, 15% test split) featuring scores across seven perceptual dimensions and temporal intervals/descriptions for noise, distortion, and unnatural pauses. Baselines include QualiSpeech-FT-SALMONN, QualiSpeech-FT (using Audio Flamingo 3 backbone), and SQ-LLM. Metrics include Pearson Correlation Coefficient (PCC) for dimension scores and MOS, F1 and Intersection over Union (IoU) for artifact localization, and ROUGE-L with GPT-4o correlation for long-form text descriptions.

Results

The proposed LLM-judge dimension-wise GRPO method achieves a new state-of-the-art mean PCC of 0.71 across dimensions and a MOS PCC of 0.76, outperforming the QualiSpeech-FT baseline (0.60 mean PCC, 0.64 MOS PCC) and SQ-LLM (0.63 mean PCC, 0.68 MOS PCC). The alternative Accuracy + Semantic Similarity reward achieves a 0.68 mean PCC and 0.72 MOS PCC.

Ablation studies reveal that unfreezing the audio encoder is vital, driving a 0.12 jump in average PCC compared to a frozen encoder. Single-stage reasoning-only models crash by up to 0.20 PCC in dimension prediction, while calibration-only models achieve high numerical scores (0.66 average PCC) but are structurally incapable of generating long-form natural language justifications and collapse on text metrics (ROUGE-L 0.12).

Systems/ConditionsMean PCCMOS PCCNoise F1Distortion IoULong-form Corr.
QualiSpeech-FT-SALMONN0.570.630.440.780.75
QualiSpeech-FT (AF3)0.600.640.560.810.78
SQ-LLM0.630.680.720.740.82
Ours (Acc. + Sem. dim.-wise)0.680.720.770.800.80
Ours (LLM-judge dim.-wise)0.710.760.770.840.83

Limitations

Unfreezing the audio encoder during calibration introduces substantial computational overhead relative to frozen encoder baselines. The fine-grained reward design is strictly bound to the predefined artifact taxonomy of the QualiSpeech benchmark, meaning performance may degrade on out-of-distribution artifacts such as ultra-low bitrate codec degradation.

Why read this

Speech and ML engineers building explainable audio evaluation models should read this to see how dimension-specific RL rewards solve the catastrophic forgetting of numerical scores during text-generation fine-tuning.

Code

Applications

Automated diagnostic evaluation of speech generation systems, telephony quality monitoring, and fine-grained acoustic artifact detection for synthetic speech verification.

Institutions

EPFL, Logitech

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2362