TL;DR — The paper introduces a multi-view noise-driven batching strategy for text-only domain adaptation in LLM-based ASR, preventing catastrophic forgetting by framing adaptation as a multi-view denoising task and achieving up to 25.4% relative WER improvement.
Key contributions
- A text-only adaptation formulation for LLM-based ASR that preserves speech-to-LLM cross-modal alignment without requiring architectural modifications or extra learnable parameters.
- A multi-view noise-driven batching scheme that combines paired source audio-text, projector-induced noisy transcripts, synthetically corrupted source transcripts, and corrupted target transcripts.
- Comprehensive evaluation across in-domain, out-of-domain, and cross-domain setups on DefinedAI and SlideSpeech datasets, yielding up to 25.4% relative WER improvements over prior text-only adaptation baselines.
- Public release of source code to facilitate reproducibility.
Problem
Adapting LLM-based automatic speech recognition systems to novel target domains typically requires expensive paired audio-text data, while relying solely on target-domain text leads to catastrophic forgetting of the cross-modal alignment established by the speech projector. Previous text-only adaptation approaches for end-to-end ASR—such as intermediate CTC adaptation, shallow fusion with external language models, text-to-mel generation, decoder adaptation for Whisper, checkpoint monitoring (Fang et al.), and soft-prompt insertion (Ma et al.)—either fail to protect the speech-text bridge or struggle under domain shifts. Solving this gap is crucial because text data is vastly more abundant and cheaper to acquire than transcribed speech.
Method
The architecture comprises a pretrained speech encoder (WavLM-Large), a learnable speech projector (single linear layer with a ReLU activation and regression layer), and a frozen pretrained LLM decoder (Llama 3.2 3B Instruct). The method models ASR inference as a denoising task where the projector maps audio into soft tokens mimicking corrupted text; text-only adaptation is achieved by substituting missing target audio with synthetic noise and fine-tuning the LLM using LoRA applied to self-attention query and value projection layers (rank 16, alpha 32, learning rate 1e-4, 1000 warm-up steps, 5 epochs).
Mini-batches are constructed using a multi-view composition strategy defined by proportions σa, σta, σt, and τ. Here, σa contains paired source audio-text (sp(a), t), σta contains projector-induced noisy transcripts from source audio (noisea(t), t), σt contains synthetically corrupted source transcripts (noise(t), t), and τ contains corrupted target-domain transcripts (noise(t), t) from Dtgt. The synthetic noise function noise(t) uses nlpaug to randomly select 15% of words and replace 30% of their characters with random symbols (1 to 10 edits per utterance), followed by a 10% chance for each character to be repeated 1 to 3 times to mimic projector duplication patterns. Proportions are set via a heuristic where τ is proportional to the relative training size of the target domain, and the remaining source shares are distributed equally as σa = σta = σt = (1 - τ)/3.
During inference, the model takes speech audio through the encoder and projector to generate the input prompt, producing the final transcription without needing any adaptation modules or target audio.
Experimental setup
Experiments use two conversational speech corpora: DefinedAI (customer-agent telephone calls, 125 total hours used across Banking, Insurance, and Healthcare splits) and SlideSpeech (YouTube conference videos, 1,000 total hours, partitioning Life, Talent, and English as source and Agriculture, Animation, and Musical Instruments as target). Baselines include the non-adapted base model, fully supervised audio-adapted models (best-case upper bound), and text-only adaptation methods by Fang et al. and Ma et al. Evaluation metric is Word Error Rate (WER, %) along with relative improvement (Δ, %).
Results
On in-domain DefinedAI tasks, the proposed method reduces WER to 6.53% on Banking (18.6% relative improvement) and 7.59% on Insurance (18.9% improvement), significantly outperforming Fang et al. (7.89% / 9.10%) and Ma et al. (7.75% / 9.24%). On out-of-domain SlideSpeech tasks, it achieves WERs of 15.56% (Ag), 15.35% (An), and 14.10% (MI), outperforming baselines across all splits. In the challenging cross-domain setup (DefinedAI source to SlideSpeech target), it records substantial relative gains, such as 25.4% improvement on Animation (25.32% WER vs 33.92% base) and 25.1% on Agriculture. Ablations confirm that omitting the source audio component (σa = 0) causes catastrophic failure, degrading WER by over 30% due to loss of speech-text alignment, and that syntactic noise perturbation outperforms random or empty text inputs.
| System | Banking (WER / ) | Insurance (WER / ) |
|---|---|---|
| Base model | 8.02 / — | 9.36 / — |
| Adapted model (audio) | 4.55 / 43.3% | 6.01 / 35.8% |
| Fang et al. [14] | 7.89 / 1.6% | 9.10 / 2.8% |
| Ma et al. [23] | 7.75 / 3.4% | 9.24 / 1.3% |
| Ours | 6.53 / 18.6% | 7.59 / 18.9% |
Limitations
The approach relies heavily on a heuristic proportion setting for τ and requires access to at least a small source-domain audio-text dataset to preserve alignment (σa > 0). While it successfully narrows the linguistic gap, it still underperforms compared to audio-based adaptation under severe cross-domain acoustic shifts. The evaluation is scoped to conversational telephone and video data in English, leaving multilingual and extreme low-resource acoustic scenarios unverified.
Why read this
Speech and ML researchers working on domain adaptation for LLM-based ASR will find a clean, parameter-efficient framework that eliminates the need for target audio by cleverly reformulating adaptation as multi-view text denoising.
Code
Applications
Adapting conversational voice assistants, customer service transcription bots, and automated meeting minutes systems to new proprietary industrial domains using unlabelled transcript text alone.
Institutions
Idiap Research Institute, EPFL, University of Zurich, Uniphore, Brno University of Technology
Funding / 經費: Idiap Research Institute and Uniphore collaboration project, EU Horizon 2020 project ELOQUENCE
Related
- Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR — same problem · relatedness 2.8/3
- Merging the Knowledge of LLMs for Automatic Speech Recognition — same problem · relatedness 2.5/3
- Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR — same problem · relatedness 2.4/3
- Weakly Masked Residual Reliability Learning for Unsupervised Domain Adaptation in Speech Models — same problem · relatedness 2.2/3
- Retention-Preserving Gradient Projection with Entropy-Guided Token-Level Distillation for Rehearsal-Free Continual ASR — same problem · relatedness 2.2/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-3422