All papers
Resources & evaluationFull-paper digest

How Frequency Band Importance Affects Neural Network Predictions and Human Perception for Speech Quality Assessment

Ada Lamba, Donald S. Williamson

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — This paper investigates what frequency bands drive speech quality assessment models (MOSNet, DNSMOS, SCOREQ) using SHAP, partial dependence plots, and perturbation analysis, and compares these findings against a human listening study. Results show that both networks and humans are substantially more sensitive to detrimental factors (like audio distortion and noise) than beneficial ones, and that top-performing models leverage broader input representations than underperforming ones.

Key contributions

  • Applies three complementary explainability methods (SHAP, PDP, and perturbation analysis) across three distinct speech quality assessment architectures (DNSMOS, MOSNet, SCOREQ).
  • Conducts a comprehensive 200-participant listening study (via Prolific) using paired preference and factor-attribution testing on perturbed speech to benchmark human perceptual sensitivity against model behavior.
  • Identifies that deep neural quality models and human listeners share an asymmetry in sensitivity, reacting much more strongly to quality-degrading artifacts than quality-enhancing factors.
  • Demonstrates that lower-performing architectures (e.g., MOSNet-CNN) restrict their focus to an overly narrow band of low frequencies, explaining their poor correlation with human evaluations.

Problem

Automated speech quality assessment models frequently achieve high internal accuracy on training sets but fail to generalize or correlate well with human mean opinion score (MOS) assessments on unseen real-world data. Prior explainability literature in speech is heavily fragmented, typically locking into a single model or a single attribution approach without cross-model validation or human perceptual grounding. Without understanding whether networks rely on the same acoustic features as humans, developers resort to blind hyperparameter tuning rather than diagnosing root architectural biases.

Method

The authors evaluate three established automated quality assessment architectures taking 161-band frequency spectrograms (50 Hz resolution): MOSNet-CNN (17-layer convolutional regression), DNSMOS (14-layer convolutional regression), and SCOREQ (a wav2vec 2.0-based transformer finetuned with triplet and contrastive losses, followed by an L2-trained linear MOS head). Local and global explanations are generated using SHAP (calculating marginal feature contributions via subset retraining), Partial Dependence Plots (PDPs, measuring expected model predictions across input feature ranges), and perturbation analysis (scaling specific frequency bands by factors of 0.25, 0.5, 10, 50, 100, and 1000 while preserving original phase).

The human listening study presented 200 participants with 16 paired audio comparisons each (original vs. perturbed by scaling one of 15 selected frequency bands at scales 0.5, 10, and 1000). Participants rated both signals on a 5-point MOS scale, stated a preference, and explicitly categorized whether factors like noise, distortion, speech-to-noise ratio (SNR), and noise type helped, hurt, or did not affect their rating.

Experimental setup

Models were evaluated on IUCOSINE, a 14,400/3,600 sample train/test split derived from the COSINE corpus containing real-world conversational speech recorded via body-worn microphones (mouth, chest, shoulder, throat) to ensure zero overlap with training sets. Baseline performances on the IUCOSINE test set measured via MSE, Linear Correlation Coefficient (LCC), and Spearman Rank Correlation Coefficient (SRCC) were: DNSMOS (MSE: 0.3890, LCC: 0.2317, SRCC: 0.1981), MOSNet (MSE: 0.5248, LCC: 0.1877, SRCC: 0.1981), and SCOREQ (MSE: 0.4624, LCC: 0.6283, SRCC: 0.6249). Experiments utilized a 48-core Dual Intel Xeon 8268s or an NVIDIA Volta V100 GPU.

Results

SCOREQ achieved the strongest alignment with human judgment, recording an LCC of 0.6283 and SRCC of 0.6249, whereas DNSMOS and MOSNet lagged significantly (LCC < 0.24). Perturbation correlations with human ratings across tested frequency bands were statistically significant and positive for SCOREQ (e.g., PCC 0.974 to 0.999 across bands 2-5) and DNSMOS, while MOSNet showed inverse or negative correlations (reaching significant negative PCC of -0.980 on bands 31 and 41). Qualitative human responses revealed that 59.8% of participants found distortion to have a detrimental effect on ratings, confirming a strong asymmetry toward penalizing negative acoustic factors.

SystemMSE (↓)LCC (↑)SRCC (↑)
DNSMOS0.38900.23170.1981
MOSNet0.52480.18770.1981
SCOREQ0.46240.62830.6249

Limitations

The perturbation analysis independently modified single frequency bands without accounting for inter-band acoustic correlations, which inadvertently introduced audible synthetic artifacts that may have biased human preference ratings. The study was constrained to English-speaking participants from the United States with normal hearing, limiting cross-linguistic and diverse demographic generalizability. Furthermore, the evaluation focused exclusively on three specific model architectures (omitting the BLSTM variant of MOSNet) and relied on a single out-of-domain evaluation corpus (IUCOSINE).

Why read this

Speech and machine learning engineers building quality assessment tools will learn how to audit black-box models using multi-faceted explainability frameworks to bridge the gap between automated metrics and human perception.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Automated speech enhancement evaluation, robust telecommunications quality monitoring, generative speech synthesis benchmarking, and acoustic model data curation.

Institutions

Ohio State University

Funding / 經費: National Science Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2382