All papers
Speech recognitionFull-paper digest

Audio-KWS-Gated Error Memory Retrieval for Incremental ASR Post-Correction

Taira Ashikawa, Daichi Hayakawa, Takehiko Kagoshima, Tomoki Watanabe

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.3 KB · Ready to paste

Preview copied content

TL;DR — The paper introduces an incremental ASR post-correction framework using open-vocabulary Audio-KWS-gated retrieval to fetch utterance-relevant historical error correction records, achieving ~70% prompt length reduction while improving WER, CER, and Bias-F1.

Key contributions

  • Proposes an incremental LLM-based ASR post-correction framework that scales past LLM context window limits by gating correction memory retrieval using Audio-KWS.
  • Derives a dynamically expanding keyword inventory from LLM error analyses, indexing canonical forms, misrecognized variants, and phonetic readings (for Japanese).
  • Demonstrates on Earnings-21 (English) and CSJ (Japanese) that ~70% prompt reduction is achievable while preserving or boosting transcription accuracy and rare/domain-specific bias-word recovery.

Problem

End-to-end ASR systems struggle with rare and domain-specific terms (named entities, technical jargon) that emerge continuously in dynamic fields like finance and news. While LLM-based second-pass post-correction and cumulative "error memory" logs can fix recurring errors without retraining, full-history prompting quickly hits LLM context window limits. Furthermore, text-driven retrieval of correction history fails when the trigger term itself is misrecognized by the ASR system, highlighting the need for audio-grounded retrieval cues.

Method

The pipeline operates incrementally across speech segments. When a human reference transcript becomes available, an LLM performs token-granularity alignment with the ASR hypothesis to extract span-level phrase pairs (hyp span ↔ ref phrase) along with error tags and edit cues. These are stored in an error memory H and indexed via an inverted index I into a dynamically expanding keyword inventory V. For Japanese, LLM-estimated phonetic readings are indexed as keywords to bypass kanji-surface mismatch issues during audio matching, while corrections apply to surface strings.

During inference, open-vocabulary Audio-KWS (AdaKWS, powered by a frozen Whisper-medium encoder and lightweight classification/conditioning modules) scores the audio against V. Keywords exceeding confidence threshold tau = 10^-4 are filtered, and the Top-N keywords are selected to retrieve associated error memory records H~t. To fit the LLM context window, retrieved records are capped at a maximum of M = 200 records by recency. The filtered memory subset alongside the ASR hypothesis hypt is fed into an LLM (openai/gpt-oss-20b) using deterministic decoding (temperature = 0.0, max completion tokens = 27000) under strict constraints to perform record-guided, conservative edits minimizing paraphrasing or deletion.

Experimental setup

Evaluated on English Earnings-21 (earnings-call transcripts, ASR hypotheses from released dataset) and Japanese CSJ (lecture spontaneous speech, hypotheses generated via Whisper-large-v3). Evaluated on bias-word subsets comprising 417 English and 442 Japanese utterances containing frequent troublesome proper nouns/technical terms. Baselines include a standard ASR output and a capped recent-history baseline without KWS (w/o KWS) limited to M = 200 records. Metrics include WER (English), CER (Japanese), micro-averaged Bias-F1, Precision, Recall, and prompt-length reduction (micro and macro percentages based on character counts). AdaKWS was trained on VoxPopuli (English) and CSJ train sets.

Results

On Earnings-21, Top-20 retrieval achieves 70.9% micro prompt reduction while improving WER from 29.06 to 28.92 and Bias-F1 from 0.836 to 0.856 compared to the w/o KWS baseline. Top-30 achieves the absolute best English WER (28.82) and Bias-F1 (0.875) with a 66.3% micro prompt reduction. On CSJ, Top-10 yields a 71.7% micro prompt reduction while improving CER from 14.86 to 14.62 and Bias-F1 from 0.648 to 0.673; Top-30 achieves the best Japanese CER (13.98) and Bias-F1 (0.728) with a 64.1% micro prompt reduction. Highly aggressive filtering (Top-1) heavily degrades WER/CER and bias recovery due to missed context.

SystemEnglish WEREnglish Bias-F1Japanese CERJapanese Bias-F1Prompt Red. (mic)
ASR (no post-correction)31.240.43516.020.548—
w/o KWS (baseline)29.060.83614.860.648—
w/ KWS Top-130.400.64215.620.59993.2%
w/ KWS Top-1029.150.86114.620.67371.7%
w/ KWS Top-2028.920.85614.090.71170.9%
w/ KWS Top-3028.820.87513.980.72866.3%

Limitations

The framework relies on the availability of human-corrected reference transcripts to asynchronously update the error memory and keyword inventory. Performance depends heavily on Audio-KWS recall; false negatives in KWS lead to missing historical correction context. The evaluation is limited to English and Japanese datasets, and end-to-end inference latency is constrained by the compute overhead of separate error analysis, KWS scanning, and LLM generation steps.

Why read this

Researchers and engineers building production-scale streaming ASR or post-correction systems should read this to see how audio-grounded retrieval effectively circumvents LLM context window bottlenecks and overcomes the text-retrieval brittleness of corrupted ASR hypotheses.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Streaming and batch conversational transcription, financial earnings call indexing, contact-center transcription post-processing, and domain-specific terminology correction.

Institutions

Toshiba

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-363