TL;DR — URGENT-MOS is a unified speech quality assessment framework that jointly models absolute multi-metric quality prediction and pairwise preference prediction under heterogeneous supervision, achieving state-of-the-art cross-domain robustness.
Key contributions
- Jointly models multi-metric absolute quality prediction and pairwise preference prediction within a single shared architecture.
- Proposes a multi-metric training strategy capable of learning from diverse datasets with incomplete/missing metric annotations using a binary validity mask.
- Derives preference-annotated training and evaluation pairs from abundant Absolute Category Rating (ACR) human MOS datasets through reference-scope, corpus-level, and arbitrary matching.
- Introduces Range-Constraining Activations (rcAct) to guarantee that predicted scores strictly respect valid metric value bounds.
Problem
Traditional speech quality assessment (SQA) frameworks optimize either absolute quality or pairwise preference independently, ignoring complementary supervision across datasets. Furthermore, annotations across public SQA corpora are highly inconsistent and heterogeneous, leading to poor cross-domain robustness when models are deployed on out-of-domain generated speech. As modern speech synthesis approaches human-level quality, subtle inter-system differences require comparison-based paradigms that directly evaluate preference rather than absolute scales alone.
Method
URGENT-MOS uses a multi-branch feature extractor where pretrained encoders (e.g., WavLM, Kimi-Audio, Qwen3OmniCaptioner, Audio Flamingo) process waveforms in parallel to produce representations . These are temporally interpolated to a common length and fused into a shared representation . The Absolute Metric Prediction Module (AMPM) predicts grouped metrics across categories like Naturalness, Intelligibility, and Speaker Similarity. The Naturalness-Conditioned Preference Module (NCPM) uses cross-attention over Naturalness-category representations of sample pairs () to predict pairwise preferences. Each metric head uses a scaled-sigmoid range-constraining activation (rcAct).
Training handles missing annotations via a binary validity mask that omits NaN labels from the loss computation, averaging losses only over valid metrics . Preference supervision is derived from ACR human MOS data using a tie threshold , augmented by reversing input pair orders during training to ensure order invariance. The system is optimized using an loss formulation.
The model is trained on 13 public and challenge datasets spanning TTS, voice conversion, speech enhancement, and telephony, totaling hundreds of thousands of samples and hundreds of hours of audio. It uses 6-layer Transformer encoders with 8 attention heads, hidden dimension 768, feed-forward dimension 2048, a constant learning rate of , and dynamic batching (~400 seconds per batch) for 60,000 steps on a single NVIDIA A100 GPU.
Experimental setup
Evaluated on diverse public and challenge test sets including SOMOS, TMHINT-QI, UR25-SQA, CHiME-7-UDASE-Eval, NISQA (FOR and LIVETALK), SpeechEval, SpeechJudge, and BC19. Baselines include single-purpose objective models like DNSMOS, UTMOS, SCOREQ, Distill-MOS, NISQA-MOS, SpeechEval, and SpeechJudge. Metrics reported are utterance-level Pearson correlation (LCC), Spearman rank correlation (SRCC), and pairwise preference accuracy ( and ).
Results
The F4C1M5 variant (4 encoders, 1 category, 5 naturalness metrics) consistently outperforms baseline models. On the SOMOS preference evaluation, F4C1M5 achieves 0.73 / 0.84 accuracy (), outperforming UTMOS (0.65 / 0.73) and DNSMOS (0.51 / 0.52). On TMHINT-QI correlation, it achieves 0.80 LCC / 0.76 SRCC, beating NISQA-MOS (0.53 / 0.35) and DNSMOS (0.41 / 0.37).
Ablations reveal that incorporating naturalness-related multi-metric supervision (M5) improves performance over MOS-only training (M1), while extending supervision to all 15 metrics (M15) degrades performance on certain datasets due to weak or inconsistent correlations among metrics like LSD and MCD.
| Model | SOMOS () | TMHINT-QI () | CHiME-7 () | LIVETALK () |
|---|---|---|---|---|
| DNSMOS | 0.51 / 0.52 | 0.63 / 0.66 | 0.51 / 0.53 | 0.71 / 0.79 |
| UTMOS | 0.65 / 0.73 | 0.69 / 0.73 | 0.57 / 0.61 | 0.80 / 0.89 |
| SpeechEval | 0.62 / 0.68 | 0.79 / 0.85 | 0.66 / 0.75 | 0.77 / 0.85 |
| F1C1M5_ | 0.71 / 0.83 | 0.79 / 0.84 | 0.70 / 0.79 | 0.76 / 0.84 |
| F4C1M5_ | 0.73 / 0.84 | 0.79 / 0.86 | 0.78 / 0.92 | 0.85 / 0.92 |
Limitations
Full multi-metric supervision (M15) causes performance degradation due to noise or weak correlations from certain low-level acoustic metrics like LSD and MCD. The approach relies heavily on pretrained speech encoders, making performance bound by the capacity and language coverage of those base models.
Why read this
Researchers building modern generative speech evaluation pipelines should read this to learn how to unify absolute metric prediction and pairwise preference learning under heterogeneous supervision.
Code
Applications
Automated evaluation of text-to-speech, voice conversion, and speech enhancement systems, as well as reward modeling for generative speech alignment.
Institutions
Shanghai Jiao Tong University, Carnegie Mellon University, Technische Universitat Braunschweig, Meta, Waseda University, VUI Labs
Funding / 經費: National Key Research and Development Program of China, China NSFC Project, SJTU Med-X Translational Research Grant
Related
- PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets — same problem · relatedness 2.8/3
- DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning — same problem · relatedness 2.8/3
- A Fine-Grained Acoustically-Aware Pre-training Encoder for Speech Quality Assessment — same problem · relatedness 2.8/3
- Calibration-Reasoning Framework for Descriptive Speech Quality Assessment — same problem · relatedness 2.6/3
- CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models — same problem · relatedness 2.5/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1671