All papers
Speech recognitionFull-paper digest

COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation

Jhih-Rong Guo, Bi-Cheng Yan, Tien-Hong Lo, Berlin Chen

Code & resourcesgithub.com/Guo0911/COALA

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.9 KB · Ready to paste

Preview copied content

TL;DR — COALA introduces a contextual biasing framework for speech-augmented language models using a discriminative scoring projector and two novel loss functions (MPD-Loss and DPD-Loss) that prevent gradient collapse in multi-entity scenarios, achieving a 99.09% Recall#20 on LibriSpeech test-clean.

Key contributions

  • Maps SLM latent representations into a specialized discriminative space via an MLP projector to score matching intensity between audio and candidate entities independently of the LLM vocabulary.
  • Proposes Multi-Positive Discriminative Loss (MPD-Loss) to eliminate inter-positive entity competition during softmax normalization in multi-target utterances.
  • Proposes Decoupled Point-wise Discriminative Loss (DPD-Loss) using a sigmoid formulation to treat entity scoring as an absolute binary classification task, avoiding relative ranking limitations.
  • Implements a Biasing Target Identification (BTI) mechanism using a top-K strategy and an <unbiased> token threshold to feed compact entity prompts into the SLM, preventing context-window OOM failures.

Problem

Integrating external knowledge via contextual biasing into Speech-augmented Language Models (SLMs) is hindered by the models' strict context-window limitations and performance degradation when processing large-scale biasing lists with multi-target utterances (where multiple rare words co-occur). Existing inference-time FST methods and training-time attention adapters suffer from error propagation or severe training collapse because standard discriminative losses induce gradient conflicts and mutual exclusivity among multiple positive entities.

Method

COALA couples a pre-trained Whisper-large-v2 audio encoder (640M parameters) with a SmolLM2-135M-Instruct backbone language model (totaling 777M parameters), incorporating a dedicated CTC module after the audio adapter to generate frame-level alignments as audio prefix tokens. For contextual biasing scoring, the last hidden states of candidate entities from the backbone LM are fed into a discriminative projector—a 2-layer multi-layer perceptron (MLP) with a ReLU activation—to compute a sequence-level, length-normalized score si for each entity ei.

To overcome the training collapse of standard discriminative loss on multi-target data (which requires auxiliary log loss), COALA trains a discriminative projector and a LoRA module using either MPD-Loss or DPD-Loss. MPD-Loss localizes the softmax normalization to a single positive entity against the entire negative set E^-, removing mutual exclusivity among co-occurring positive targets. DPD-Loss goes further by leveraging the sigmoid function sigma(·) to treat each entity independently as a point-wise binary classification task, optimizing absolute matching intensities and establishing a stable decision boundary centered around zero.

During inference, a Biasing Target Identification (BTI) stage ranks candidate entities and selects a top-K (K=10) subset using the score of a special <unbiased> token as a dynamic threshold to filter low-scoring entities, ensuring the SLM prompt stays within length limits while feeding precisely targeted rare words.

Experimental setup

Evaluated on the LibriSpeech corpus (960-hour full training set, dev-clean validation, and test-clean/test-other test sets) with biasing list sizes N = 500, 1000, 5000 containing common words (5K high frequency) and rare words (209.2K low frequency). Compared against baseline systems including CTC-Filter, K-Prompt, and standard Bias-Loss (discriminative loss + log loss), measured using Word Error Rate (WER, decomposed into U-WER and B-WER) and Recall metrics (Recall@X, Recall#X). Implemented using a 24GB RTX 3090 GPU in a two-stage recipe: Stage 1 trains the audio adapter and CTC module while tuning backbone LoRA for 10 epochs (batch size 4); Stage 2 freezes previous layers and trains the discriminative projector and a new LoRA for 7 epochs (batch size 1, M=120 candidates).

Results

On the biasing scoring task with N = 5000, DPD-Loss achieves a Recall#20 of 99.09% on test-clean and 96.59% on test-other, outperforming the baseline Bias-Loss (98.72% and 95.10%) and prompt-based methods like CTC-Filter (93.26% and 83.83%). For downstream contextual ASR, integrating DPD-Loss with the BTI mechanism drops B-WER on test-clean down to 3.25% (N=500) and 3.86% (N=1000), compared to 23.39% without biasing. Unfiltered large biasing lists (N=5000) fail due to out-of-memory (OOM) errors on the 24GB VRAM hardware, highlighting the necessity of the proposed BTI module.

MethodsRecall#20 clean (%)Recall#20 other (%)Recall#50 clean (%)Recall#50 other (%)
CTC-Filter93.2683.8394.4585.49
K-Prompt86.3073.8888.9279.05
Bias-Loss [13]98.7295.1099.2997.27
MPD-Loss (Our)87.8695.6599.4397.55
DPD-Loss (Our)99.0996.5999.6098.17

Limitations

The framework relies on a top-K strategy (K=10) that structurally caps retrieval performance on utterances containing more than 11 target entities. Unfiltered large-scale biasing lists (N=5000) cause OOM errors on standard 24GB GPUs without the BTI pruning mechanism. Evaluation is restricted to the English LibriSpeech corpus, leaving multilingual and domain-transfer scalability unverified.

Why read this

Speech and ML researchers working on E2E contextual biasing should read this paper to understand how formulating entity scoring as a point-wise binary classification task (DPD-Loss) eliminates gradient conflicts in multi-target speech language models.

Code

Applications

Voice assistants, command-and-control systems, and domain-specific speech recognition (e.g., medical, legal, or contact-list dialing) requiring robust recognition of rare or proprietary entity names.

Institutions

National Taiwan Normal University

Funding / 經費: Realtek Semiconductor Corporation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1097