---
id: tao26_interspeech
title: "ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint
  Multi-Resolution Speech Quality Modeling"
authors:
  - Zhuoyan Tao
  - Jiatong Shi
  - Hye-jin Shim
  - Shinji Watanabe
year: 2026
doi: 10.21437/Interspeech.2026-927
isca_url: https://www.isca-archive.org/interspeech_2026/tao26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/tao26_interspeech.pdf
session: Speech and Audio Quality Assessment
topics:
  - speech-enhancement
  - self-supervised
  - evaluation
category: resources-evaluation
labels:
  - streaming-real-time
institutions:
  - University of Southern California
  - Carnegie Mellon University
funding:
  - "Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support"
  - National Science Foundation
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: tao26_interspeech
  category: resources-evaluation
  labels:
    - streaming-real-time
  institutions:
    - University of Southern California
    - Carnegie Mellon University
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-927
  pdf: https://www.isca-archive.org/interspeech_2026/tao26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/tao26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/tao26_interspeech/markdown.md
---

# ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling

*Zhuoyan Tao, Jiatong Shi, Hye-jin Shim, Shinji Watanabe*

[PDF](https://www.isca-archive.org/interspeech_2026/tao26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/tao26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-927)

**Category:** `resources-evaluation` · **Labels:** `streaming-real-time`

**TL;DR** — ANCHOR reformulates speech quality estimation as a multi-resolution autoregressive prediction task to evaluate partial audio prefixes accurately, achieving a 48% reduction in PLCMOS error on 2-second prefixes. It introduces a resolution-aware decoding hierarchy that generates chunk-level scores before full-utterance scores within a unified sequence.

## Key contributions

- Joint chunk-level and full-utterance multi-metric supervision within a unified autoregressive framework.
- A resolution-aware decoding hierarchy enforcing a chunk-first coarse-to-fine prediction schedule.
- Prefix-to-full convergence analysis revealing an effective perceptual context horizon of 4-6 seconds.
- A controlled distortion stress test isolating structured extrapolation biases under localized corruption.

## Problem

Most existing objective metrics like PESQ, ViSQOL, and learned non-intrusive predictors like UTMOS and DNSMOS assume full-context availability, evaluating signals only after complete utterances are observed. This creates a severe mismatch for streaming applications, generative speech models, and systems dealing with short packet-loss bursts or clipping events where early, prefix-constrained quality estimation is necessary. Global context pooling in standard models smooths out temporally sparse distortions and fails to reflect local perceptual degradation until much later.

## Method

ANCHOR builds upon the ARECHO architecture, using a frozen WavLM-Large acoustic frontend paired with a 4-layer audio encoder and a 12-layer Transformer decoder (8 attention heads, embedding dimension 256). It expands the decoder vocabulary from 32,926 to 65,828 tokens to support dual-resolution query tokens. Continuous metrics are discretized into 500 percentile-based bins (with signed log compression applied to heavy-tailed metrics like SI-SNR).

The model enforces a chunk-first decoding order where chunk-level target tokens are generated before full-utterance tokens, effectively using local quality estimates as intermediate conditional latents. Training uses the Overall Base configuration dataset (308.8 hours across 170,013 utterances) expanded via cumulative prefixes at 2, 4, 6, and 8 seconds, yielding 583,983 samples (467,657 training / 116,326 validation instances). It is optimized using the AdamW optimizer with a learning rate of 4e-4, linear warmup over 50k steps, 0.1 label smoothing, batch size 12, gradient accumulation of 2, and trained for 15 epochs from a pretrained ARECHO checkpoint.

## Experimental setup

Experiments use the Overall Base dataset (308.8 hours, 170,013 utterances) expanded to 583,983 prefix samples, evaluated on the Overall Dev split (34,726 prefix instances after expansion). The primary baseline is the pretrained ARECHO checkpoint applied directly to prefix inputs without adaptation. Evaluation metrics include Mean Absolute Error (MAE), Pearson Correlation Coefficient (LCC/PCC), and Spearman Rank Correlation (SRCC) for both chunk-level and full-utterance tasks.

## Results

On chunk-level prediction, ANCHOR achieves a 48% MAE reduction for PLCMOS on 2-second prefixes (maintaining consistent gains: 33% at 4s, 16% at 6s, 12% at 8s). For UTMOS, ANCHOR improves MAE at 2s (0.241 to 0.214; PCC 0.935 to 0.950), though ARECHO surpasses it at longer prefixes due to the local attention shift. For full-utterance prediction from prefixes, the largest MAE drop occurs between 2s and 4s across all metrics, with Pearson correlation stabilizing by 4-6 seconds, defining an effective context horizon.

| Metric & Prefix Length | ANCHOR MAE | ARECHO MAE | ANCHOR LCC | ARECHO LCC |
|---|---|---|---|---|
| PLCMOS (2s) | 0.865 | - | 0.629 | - |
| PLCMOS (4s) | 0.725 | - | 0.719 | - |
| UTMOS (2s) | 0.236 | - | 0.934 | - |
| UTMOS (4s) | 0.183 | - | 0.959 | - |
| DNS (4s) | 0.238 | - | 0.895 | - |

## Limitations

ANCHOR currently relies on a non-formal streaming frontend (specifically non-causal WavLM-Large) and therefore does not constitute a fully causal streaming system. Formal component-wise ablations comparing interleaved versus chunk-first decoding orders were omitted. The evaluation is bound to datasets spanning 308.8 hours and prefix lengths between 2 and 8 seconds.

## Why read this

Speech and ML engineers building streaming communication or autoregressive generative speech models should read this to learn how hierarchical, chunk-first decoding can resolve prefix-constrained quality estimation trade-offs.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Real-time streaming communication quality monitoring, generative speech model evaluation, and low-latency audio packet-loss tracking.

## Institutions / 機構

University of Southern California, Carnegie Mellon University

**Funding / 經費:** Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support, National Science Foundation

## Related

- [DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning](liang26_interspeech.md) — same problem · relatedness 2.4/3
- [A Fine-Grained Acoustically-Aware Pre-training Encoder for Speech Quality Assessment](sultana26_interspeech.md) — same problem · relatedness 2.3/3
- [Calibration-Reasoning Framework for Descriptive Speech Quality Assessment](kostenok26_interspeech.md) — same problem · relatedness 2.3/3
- [CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models](ferreira26_interspeech.md) — same problem · relatedness 2.2/3
- [PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets](fan26_interspeech.md) — same problem · relatedness 2.1/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
