All papers
Speech recognitionFull-paper digest

Rank-Distance Based Confidence Estimation for ASR

Nagarathna Ravi, Madduri Aiswarya Lakshmi, Ragesh M, Rajalakshmi Elangovan

Code & resourcesgithub.com/Nagarathna-R/2026_RanDiS_Interspeech

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.5 KB · Ready to paste

Preview copied content

TL;DR — The paper introduces RanD, a continuous target score for ASR confidence estimation that combines probability rank and distribution distance to prevent the probability-collapse issues seen in prior true-class methods. Evaluated across CTC, RNN-T, TDT, and AED architectures on Hindi and English datasets, RanD-CEM consistently outperforms state-of-the-art baselines under both in-domain and out-of-domain evaluation conditions.

Key contributions

  • Identifies and explains the true-class probability collapse phenomenon in deep ASR decoders, where incorrect token probabilities diminish toward zero as vocabulary size grows.
  • Proposes RanD, a novel continuous confidence score combining a normalized rank score (sranks_{\text{rank}}) of the true class and a normalized Euclidean distance (sdiss_{\text{dis}}) between predicted and actual posterior distributions.
  • Implements and evaluates RanD-CEM across four distinct ASR backends: CTC, RNN-T, TDT, and AED, showcasing a unified confidence estimation recipe.
  • Demonstrates robust generalizability on mismatched/out-of-domain evaluation data (e.g., NPTEL, Svarah, PB Hindi) without architecture-specific manual temporal alignments.

Problem

Maximum class probabilities and entropy transformations fail to yield reliable confidence scores because overconfident deep ASR decoders output highly skewed distributions. Existing trainable auxiliary Confidence Estimation Models (CEMs) either rely on binary targets—which completely miss partial correctness (treating substituted/inserted tokens uniformly as zero)—or continuous targets like TeLeS and TruCLeS that suffer from brittle temporal alignment propagation or probability collapse in large vocabularies. This lack of robust calibration hinders safe error correction and downstream cascading systems.

Method

The framework extracts word-level representations from trained ASR models and aligns hypothesis and reference transcripts via edit-distance. For each predicted token cj′c'_j aligned with reference cjc_j, the posterior probability vector pj′\mathbf{p}_{j'} and one-hot target vector qj′\mathbf{q}_{j'} are analyzed. The normalized rank score is computed as srank=1−r(cj)−1∣C∣−1s_{\text{rank}} = 1 - \frac{r(c_j)-1}{|C|-1}, and the normalized distribution distance score as sdis=1−∥pj′−qj′∥22s_{\text{dis}} = 1 - \frac{\|\mathbf{p}_{j'} - \mathbf{q}_{j'}\|_2}{\sqrt{2}}. The final token-level score blends them with α=0.5\alpha=0.5: s=αsrank+(1−α)sdiss = \alpha s_{\text{rank}} + (1-\alpha)s_{\text{dis}}, and word-level scores (sw′s_{w'}) average the underlying token scores.

For CTC, the CEM takes encoder hidden states, decoder softmax outputs, and posteriors and uses fully connected layers (512-256-128 neurons) with ReLU. For RNN-T and TDT, the auxiliary model uses two bidirectional LSTM layers (512 hidden units) followed by a 1024-neuron fully connected layer. For AED, it uses a 256-neuron feedforward network. All CEM architectures are trained for 50 epochs using the Adam optimizer (learning rate 10−410^{-4}), shrinkage loss, and time/frequency masking augmentation.

Experimental setup

Evaluated using pre-trained models from the NeMo framework: Hindi Conformer-CTC (trained on KB Hindi, tested on PB Hindi), Conformer-RNN-T, Parakeet-TDT, and Canary-Flash AED (the latter three trained on LibriSpeech and tested on NPTEL and Svarah Indian English datasets). Performance is measured using MAE, KLD, JSD, NCE, ECE, AUROC, and AUPRC.

Results

RanD-CEM consistently outperforms MCP, entropy-baseline, binary CEM, and prior continuous targets (TeLeS/TruCLeS) across most metrics and domains. For instance, on CTC-ASR with the KB dataset, RanD achieves an MAE of 0.0570 and NCE of 0.3886, compared to TruCLeS (MAE 0.0870, NCE 0.2971) and MCP (MAE 0.1343). On challenging mismatched domains like NPTEL and Svarah, RanD preserves low MAE and high AUPRC (e.g., reaching 0.9355 AUPRC on NPTEL for RNN-T), whereas baseline methods experience sharp degradations.

System / ConditionMAE (↓)KLD (↓)JSD (↓)NCE (↑)ECE (↓)AUROC (↑)
MCP (CTC / KB)0.13430.49950.2792-0.28230.27030.7616
TeLeS (CTC / KB)0.10780.14980.04300.14080.05240.8155
TruCLeS (CTC / KB)0.08700.10930.02780.29710.01070.8578
RanD (CTC / KB)0.05700.10280.01810.38860.04730.8669

Limitations

The current formulation requires access to rich internal token posterior distributions and intermediate representations from the parent ASR, limiting its plug-and-play usage on black-box or API-based speech models. Additionally, the approach currently struggles to natively predict completely deleted words (insertions/substitutions are modeled, but deletions require context-window expansion around gaps).

Why read this

Researchers and engineers building confidence-aware speech transcription pipelines or downstream error-correction modules should read this paper to adopt a robust, architecture-agnostic continuous confidence scoring function that bypasses fragile temporal alignments and overconfidence collapse.

Code

Applications

Automated ASR error correction, selective downstream ingestion for speech translation, disfluency detection, and voice assistant safety filtering.

Institutions

CSIR Fourth Paradigm Institute

Funding / 經費: Council of Scientific and Industrial Research, Anusandhan National Research Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1355