All papers
Resources & evaluationFull-paper digest

ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling

Zhuoyan Tao, Jiatong Shi, Hye-jin Shim, Shinji Watanabe

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.1 KB · Ready to paste

Preview copied content

TL;DR — ANCHOR reformulates speech quality estimation as a multi-resolution autoregressive prediction task to evaluate partial audio prefixes accurately, achieving a 48% reduction in PLCMOS error on 2-second prefixes. It introduces a resolution-aware decoding hierarchy that generates chunk-level scores before full-utterance scores within a unified sequence.

Key contributions

  • Joint chunk-level and full-utterance multi-metric supervision within a unified autoregressive framework.
  • A resolution-aware decoding hierarchy enforcing a chunk-first coarse-to-fine prediction schedule.
  • Prefix-to-full convergence analysis revealing an effective perceptual context horizon of 4-6 seconds.
  • A controlled distortion stress test isolating structured extrapolation biases under localized corruption.

Problem

Most existing objective metrics like PESQ, ViSQOL, and learned non-intrusive predictors like UTMOS and DNSMOS assume full-context availability, evaluating signals only after complete utterances are observed. This creates a severe mismatch for streaming applications, generative speech models, and systems dealing with short packet-loss bursts or clipping events where early, prefix-constrained quality estimation is necessary. Global context pooling in standard models smooths out temporally sparse distortions and fails to reflect local perceptual degradation until much later.

Method

ANCHOR builds upon the ARECHO architecture, using a frozen WavLM-Large acoustic frontend paired with a 4-layer audio encoder and a 12-layer Transformer decoder (8 attention heads, embedding dimension 256). It expands the decoder vocabulary from 32,926 to 65,828 tokens to support dual-resolution query tokens. Continuous metrics are discretized into 500 percentile-based bins (with signed log compression applied to heavy-tailed metrics like SI-SNR).

The model enforces a chunk-first decoding order where chunk-level target tokens are generated before full-utterance tokens, effectively using local quality estimates as intermediate conditional latents. Training uses the Overall Base configuration dataset (308.8 hours across 170,013 utterances) expanded via cumulative prefixes at 2, 4, 6, and 8 seconds, yielding 583,983 samples (467,657 training / 116,326 validation instances). It is optimized using the AdamW optimizer with a learning rate of 4e-4, linear warmup over 50k steps, 0.1 label smoothing, batch size 12, gradient accumulation of 2, and trained for 15 epochs from a pretrained ARECHO checkpoint.

Experimental setup

Experiments use the Overall Base dataset (308.8 hours, 170,013 utterances) expanded to 583,983 prefix samples, evaluated on the Overall Dev split (34,726 prefix instances after expansion). The primary baseline is the pretrained ARECHO checkpoint applied directly to prefix inputs without adaptation. Evaluation metrics include Mean Absolute Error (MAE), Pearson Correlation Coefficient (LCC/PCC), and Spearman Rank Correlation (SRCC) for both chunk-level and full-utterance tasks.

Results

On chunk-level prediction, ANCHOR achieves a 48% MAE reduction for PLCMOS on 2-second prefixes (maintaining consistent gains: 33% at 4s, 16% at 6s, 12% at 8s). For UTMOS, ANCHOR improves MAE at 2s (0.241 to 0.214; PCC 0.935 to 0.950), though ARECHO surpasses it at longer prefixes due to the local attention shift. For full-utterance prediction from prefixes, the largest MAE drop occurs between 2s and 4s across all metrics, with Pearson correlation stabilizing by 4-6 seconds, defining an effective context horizon.

Metric & Prefix LengthANCHOR MAEARECHO MAEANCHOR LCCARECHO LCC
PLCMOS (2s)0.865-0.629-
PLCMOS (4s)0.725-0.719-
UTMOS (2s)0.236-0.934-
UTMOS (4s)0.183-0.959-
DNS (4s)0.238-0.895-

Limitations

ANCHOR currently relies on a non-formal streaming frontend (specifically non-causal WavLM-Large) and therefore does not constitute a fully causal streaming system. Formal component-wise ablations comparing interleaved versus chunk-first decoding orders were omitted. The evaluation is bound to datasets spanning 308.8 hours and prefix lengths between 2 and 8 seconds.

Why read this

Speech and ML engineers building streaming communication or autoregressive generative speech models should read this to learn how hierarchical, chunk-first decoding can resolve prefix-constrained quality estimation trade-offs.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Real-time streaming communication quality monitoring, generative speech model evaluation, and low-latency audio packet-loss tracking.

Institutions

University of Southern California, Carnegie Mellon University

Funding / 經費: Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support, National Science Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-927