---
id: li26d_interspeech
title: "CAQA-Net: Continual Audio Quality Assessment Across Speech and Music Domains"
authors:
  - Naiyuan Li
  - Xiaoxun Wu
  - Yuheng Huang
  - Diqun Yan
year: 2026
doi: 10.21437/Interspeech.2026-476
isca_url: https://www.isca-archive.org/interspeech_2026/li26d_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/li26d_interspeech.pdf
session: Evaluation, Benchmarking, and Reliability of Audio Systems
topics:
  - speech-enhancement
  - self-supervised
  - evaluation
category: resources-evaluation
institutions:
  - Ningbo University
  - Ningbo University of Finance and Economics
funding:
  - National Natural Science Foundation of China
  - Zhejiang Provincial Collaborative Innovation Center for Digital Supply Chain
    and Artificial Intelligence of Bulk Commodities
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: li26d_interspeech
  category: resources-evaluation
  institutions:
    - Ningbo University
    - Ningbo University of Finance and Economics
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-476
  pdf: https://www.isca-archive.org/interspeech_2026/li26d_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/li26d_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/li26d_interspeech/markdown.md
---

# CAQA-Net: Continual Audio Quality Assessment Across Speech and Music Domains

*Naiyuan Li, Xiaoxun Wu, Yuheng Huang, Diqun Yan*

[PDF](https://www.isca-archive.org/interspeech_2026/li26d_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/li26d_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-476)

**Category:** `resources-evaluation`

**TL;DR** — CAQA-Net is a continual learning framework for audio quality assessment that handles sequential speech and music tasks without catastrophic forgetting, achieving a mean Spearman rank correlation coefficient (mSRCC) of 0.754—within 4.6% of the joint-learning upper bound.

## Key contributions

- Formulates a systematic continual learning benchmark spanning five speech and music datasets with evolving synthetic and realistic distortions.
- Proposes a dual-branch architecture combining a trainable M2D waveform encoder for plasticity and a frozen BEATS spectrogram encoder as a semantic stability anchor.
- Introduces a ranking-based loss function coupled with Learning without Forgetting (LwF) style output-level distillation to handle noisy subjective MOS annotations.
- Implements a prototype-based softmin gating mechanism on frozen anchor features enabling task-agnostic inference without knowing sample task identities.

## Problem

Static audio quality assessment models fail to keep pace with dynamic real-world audio data generated by emerging speech synthesis, voice conversion, and music generation systems. While joint learning retrains models on all available data from scratch, it incurs prohibitive computational and storage overhead, whereas naive fine-tuning causes severe catastrophic forgetting. Furthermore, continual learning in audio must overcome large acoustic domain shifts from speech to music, continuous perceptual score regression rather than discrete classification, and inconsistent scoring standards or label noise across datasets.

## Method

CAQA-Net utilizes a dual-branch feature extractor feeding independent linear prediction heads ($H_t$) for each task $t$ to decouple shared feature extraction from task-specific decisions. The M2D branch processes raw waveforms and is fine-tuned to ensure plasticity and adapt to distribution shifts, while the BEATS branch processes log-mel spectrograms, remains completely frozen, and acts as a task-invariant semantic anchor. Features from both branches are concatenated and passed to the active prediction head.

To handle subjective MOS scale variability and annotator noise, the training objective uses a pairwise relative ranking loss based on Thurstone's Case V model, mapping score differences to probabilities via an indicator function and a small numerical stability constant. To prevent catastrophic forgetting, an output-level knowledge-distillation regularizer (inspired by LwF) constrains output probability distributions rather than imposing rigid parameter-importance penalties like EWC or MAS, allowing the waveform encoder to flexibly absorb large domain shifts.

During inference, task-agnostic routing is achieved via a prototype-based adaptive gating mechanism. K-means clustering ($N=64$) is applied to BEATS features to extract task prototypes, and test samples are assigned head weights via a softmin operator over minimum squared Euclidean distances to these prototypes with a temperature parameter $\alpha=16$.

## Experimental setup

Evaluated on a 5-task sequence: TCD-VOIP (384 samples, speech), NISQA-SIM (12.5k samples, speech), Tencent Corpus (14k samples, speech), SingMOS (3.2k samples, music), and MusicEval (1.9k samples, music). Datasets are split 80%/10%/10% for train/validation/test. Optimized using Adam with learning rate $2\times 10^{-4}$ and batch size 16 for 10 epochs per task (3 warm-up epochs for the prediction head, followed by 7 epochs of joint training with early stopping). Baselines include Separate Learning (SL), Joint Learning (JL), Fine-Tune (FT), Multi-Head Fine-Tune (MH-FT), and regularization variants CAQA-Net (EWC, $\lambda=1000$), CAQA-Net (MAS, $\lambda=100$), and CAQA-Net (LwF, $\lambda=100$). Evaluation metrics include mSRCC, mean Plasticity Index (mPI), mean Stability Index (mSI), and mean Plasticity-Stability Index (mPSI).

## Results

CAQA-Net (LwF) achieves an mSRCC of 0.754, mPI of 0.862, mSI of 0.911, and mPSI of 0.887, significantly outperforming MH-FT (mSRCC 0.589) and outclassing alternative continual learning regularizers like EWC (0.590) and MAS (0.625), coming within 4.6% of the joint-learning upper bound (0.790). Ablations confirm that fusing a trainable M2D branch with a frozen BEATS branch yields the optimal balance (mPSI 0.903), whereas making BEATS trainable collapses plasticity (mPI 0.587). The prototype-based gating mechanism (mSRCC 0.754) closely approaches the performance of the oracle task-specific head (mSRCC 0.788). The framework exhibits strong order-robustness across five distinct task sequences, though frequent domain oscillations between speech and music slightly depress stability.

| Method | mSRCC | mPI | mSI | mPSI |
|---|---|---|---|---|
| SL | 0.876 | / | / | / |
| JL | 0.790 | / | / | / |
| FT | 0.541 | 0.767 | 0.661 | 0.714 |
| MH-FT | 0.589 | 0.791 | 0.709 | 0.750 |
| CAQA-Net (EWC) | 0.590 | 0.808 | 0.867 | 0.838 |
| CAQA-Net (LwF) | 0.754 | 0.862 | 0.911 | 0.887 |

## Limitations

The framework suffers minor instability when exposed to highly frequent domain oscillations between speech and music tasks. Evaluation is limited to five datasets covering synthetic and algorithmic distortions, leaving real-world unconstrained recording degradations and a wider array of generative audio domains unexplored.

## Why read this

Speech and ML engineers building streaming or evolving audio quality assessment systems will find a blueprint for balancing stability and plasticity without full network retraining. Researchers gain empirical insight into why output-level distillation outperforms parameter-level regularizers for continuous MOS regression.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Automated monitoring of streaming audio, text-to-speech generation pipelines, voice conversion systems, and online meeting platform quality control.

## Institutions / 機構

Ningbo University, Ningbo University of Finance and Economics

**Funding / 經費:** National Natural Science Foundation of China, Zhejiang Provincial Collaborative Innovation Center for Digital Supply Chain and Artificial Intelligence of Bulk Commodities

## Related

- [Calibration-Reasoning Framework for Descriptive Speech Quality Assessment](kostenok26_interspeech.md) — same problem · relatedness 2.1/3
- [DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning](liang26_interspeech.md) — same problem · relatedness 2.0/3
- [Evaluating Objective Speech Quality Metrics for Neural Audio Codecs](lanzendoerfer26_interspeech.md) — same problem · relatedness 2.0/3
- [CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models](ferreira26_interspeech.md) — same problem · relatedness 2.0/3
- [Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification](mylvaganam26_interspeech.md) — shared technique · relatedness 1.9/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
