All papers
Speech recognitionFull-paper digest

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

Zhihan Li, Hankun Wang, Yiwei Guo, Bohan Li, Xie Chen, Kai Yu

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.5 KB · Ready to paste

Preview copied content

TL;DR — READ (Reference-free Hypothesis Evaluation with Acoustic Discrepancy) evaluates ASR hypotheses directly from speech audio using an off-the-shelf autoregressive TTS model, achieving up to 21.46% relative error reduction via n-best rescoring and segment-level combination.

Key contributions

  • Proposes a training-free, reference-free ASR hypothesis evaluation metric (READ) based on conditional negative log-likelihood derived from a pre-trained autoregressive TTS model.
  • Leverages internal cross-attention maps from the TTS decoder to establish a monotonic frame-to-token alignment without requiring external aligners.
  • Enables fine-grained, segment-level system combination and n-best rescoring by exploiting the locality of acoustic discrepancy spikes.
  • Demonstrates robust performance improvements, particularly under challenging high-noise and code-switching acoustic conditions.

Problem

Traditional ASR evaluation relies heavily on reference transcripts like Word Error Rate (WER), which are unavailable in large-scale unsupervised or self-improvement settings. Existing reference-free methods—such as internal ASR confidence scores (prone to overconfidence and poor calibration), external text-only language model rescoring (which completely ignores the acoustic signal), and supervised Quality Estimation models (requiring error-labeled training data)—fail to properly balance linguistic plausibility with fine-grained acoustic grounding. Without explicit acoustic evaluation, current systems lack interpretable, diagnostic feedback to localize errors and refine hypotheses directly from raw speech signals.

Method

READ evaluates a text hypothesis XX given a speech token sequence YY using a pre-trained discrete autoregressive TTS model (CosyVoice2) in teacher-forcing mode. It computes the conditional negative log-likelihood (NLL) at each frame tt as READt=−log⁡Ptheta(yt∣x,y<t)\text{READ}_t = -\log P_theta(y_t | x, y_{<t}), serving as a fine-grained acoustic discrepancy map where loss spikes indicate misaligned or erroneous text segments. To map these frame-level scores back to text tokens, READ extracts the T×NT \times N self-attention submatrix from the TTS decoder layers and solves for a monotonic mapping π∗\pi^* via dynamic programming, maintaining internal consistency without relying on external alignment models.

For hypothesis refinement, READ is applied in three ways: (1) Sentence-level rescoring, where the total negative log-likelihood of the speech under different candidate hypotheses is summed to rank nn-best lists (with a mild 0.95 scaling factor applied to retain base language model priors); (2) Segment-level combination, which divides audio into alternating consensus and disputed intervals across multiple ASR hypotheses, selecting the segment variant that minimizes READ; and (3) ROVER integration, where segment-level combinations are fed alongside original candidates into ROVER to inject regional acoustic bias.

Experimental setup

Evaluated on LibriSpeech (test-clean, test-other), SPGISpeech, Switchboard, TEDLIUM3, VCTK-noisy, and Mandarin-English code-switching sets (ASRU2019, TALCS). Noise-augmented sets are constructed using WHAM! test noise at 0, 10, and 20 dB SNR. Uses CosyVoice2 official checkpoints without task-specific fine-tuning. Baselines include single ASR models (Whisper medium/large-v3, NVIDIA NeMo, Qwen2.5-Omni) and ROVER system combination.

Results

On n-best rescoring over Whisper-large-v3 candidates, READ reduced WER by 7.28% on LibriSpeech-clean, 20.92% on TALCS-test, 20.57% on Switchboard-test, and 21.46% on SPGISpeech-val. In segment-level combination across four disparate ASR systems, READ consistently improved performance over single bests (e.g., VCTK-noisy dropping from a best single of 2.85% to 1.84% with segment combination). The evaluation metric demonstrated stronger correlation with WER as noise increased (Pearson correlation rising from 0.7213 on clean speech to 0.8539 at 0 dB SNR). Limitations include occasional performance degradation when disputed intervals are overly long, causing segment combination to degenerate into sentence-level selection.

DatasetWhisper large-v3Whisper mediumNeMoQwen2.5-OmniROVEROurs (Segment)
LS-clean2.202.791.671.741.511.67
LS-other4.167.523.653.453.183.39
VCTK-noisy8.8718.332.852.471.842.18
ASRU-test10.3511.9821.708.009.047.60
TALCS-test16.7720.7444.479.2120.849.61

Limitations

The approach assumes that the underlying TTS model's acoustic space aligns well with diverse acoustic environments, though extremely distorted audio can degrade TTS likelihood estimation. The segment-level combination strategy relies on a greedy selection scheme and a locality assumption that can fail if disputed intervals are excessively long. Furthermore, READ currently evaluates total acoustic discrepancy but requires further exploration to explicitly disentangle substitution, deletion, and insertion error types.

Why read this

Speech and ML researchers working on unsupervised ASR improvement, confidence calibration, or system combination should read this to learn how to repurpose off-the-shelf autoregressive TTS models as powerful acoustic critics without any supervised training.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Unsupervised ASR hypothesis selection, n-best list rescoring, robust multi-system combination under noisy conditions, and error localization.

Institutions

Shanghai Jiao Tong University

Funding / 經費: China NSFC Project

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3434