All papers
Speech recognitionFull-paper digest

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

Thibault Bañeras-Roux, Sergio Burdisso, Esaú Villatoro-Tello, Dairazalia Sánchez-Cortés, Shiran Liu, Severin Baroudi, Shashi Kumar, Hasindri Watawana, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke

Code & resourcesgithub.com/idiap/llm-asr-text-only-adaptation

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

14.3 KB · Ready to paste

Preview copied content

TL;DR — This paper investigates domain adaptation for LLM-based ASR when target-domain audio is scarce but text is abundant, proposing a mixed batching strategy that combines text-only data with a small fraction of speech. Using only 10% target-domain speech (under 4 hours) in mixed batches achieves word error rates comparable to or better than conventional fine-tuning with 100% of the speech data.

Key contributions

  • Analyzed the trade-off between target-domain specialization and source-domain retention as a function of target data proportions in LLM-based ASR batch composition.
  • Proposed a hybrid mixed-batch adaptation method integrating limited paired speech-text data into a text-heavy adaptation loop to bridge the cross-modal alignment gap.
  • Demonstrated that the mixed batching strategy reduces catastrophic forgetting on source domains while boosting target-domain recognition compared to standard ASR fine-tuning.
  • Released open-source code for reproducibility of the proposed domain adaptation pipelines.

Problem

Conventional end-to-end and LLM-based automatic speech recognition systems rely heavily on large amounts of paired speech-text data for domain adaptation, which is expensive and scarce in specialized verticals like healthcare, finance, or agriculture. While adapting LLMs using abundant in-domain text-only data is a scalable alternative, it introduces a severe modality gap because the LLM unlearns the noisy acoustic representations generated by upstream speech encoders. Prior text-only adaptation methods rely on soft prompting or decoding tweaks, failing to evaluate whether minimal amounts of target-domain audio can preserve cross-modal alignment without requiring full speech collection.

Method

The architecture follows the SLAM-ASR framework, comprising a frozen WavLM-Large speech encoder, a linear speech projector downsampling acoustic frames by factor k, and a Llama-3.2-3B-Instruct LLM. The baseline model is first pretrained on source-domain paired speech-text data from DefinedAI (Banking, Insurance, and Healthcare partitions, totaling 38h 10m) where only the speech projector is fine-tuned while the encoder remains frozen. Domain adaptation is then performed by updating only LoRA adapters inserted into the LLM for 5 epochs with a batch size of 10 under a fixed instruction prompt template.

The mixed batching (MB) strategy constructs training batches by dividing data into source and target components. Source data comprises three segments: (1) paired source speech-text (σa\sigma_a), (2) projector-induced noisy tokens mapped from speech projections to nearest LLM embeddings (σta\sigma_{ta}), and (3) synthetically corrupted source transcripts generated via random character-level perturbations using nlpaug (σt\sigma_t). The target portion consists of: (1) paired target speech-text (τa\tau_a), and (2) synthetically corrupted target transcripts (τt\tau_t) using identical character perturbations. The total target proportion τ=τt+τa\tau = \tau_t + \tau_a is set to 50% based on sweep experiments showing optimal performance at this threshold, balancing underfitting/catastrophic forgetting against overfitting.

Experimental setup

Experiments use the DefinedAI corpus for in-domain evaluation (Banking partition: 36h 20m train, 2h 08m dev, 4h 16m test) and SlideSpeech for out-of-distribution evaluation (Agriculture: 29h 20m train; Musical Instruments: 8h 30m train). Evaluations compare base models, text-only adaptation, standard ASR fine-tuning (paired speech-text only), and mixed-batch adaptation across varying fractions of target audio. The primary evaluation metric is Word Error Rate (WER). Models use Llama-3.2-3B-Instruct and WavLM-Large, trained for 5 epochs with batch size 10.

Results

On the DefinedAI Banking in-domain test set, standard ASR adaptation yields a WER of 4.55% when using 100% of the speech data, whereas text-only adaptation reaches 6.38%, demonstrating a 1.83% absolute WER penalty (approx. 29% relative degradation) due to the modality gap. However, mixed-batch adaptation achieves competitive or superior performance using only 10% of the target speech data (approx. 3.5 hours for Banking, yielding WERs around 4.19%). For out-of-domain evaluation on SlideSpeech Agriculture, mixed-batch adaptation with 10% target audio matches or beats the 100% audio baseline (15.91% vs 16.04% WER). On Musical Instruments, mixed batching yields 16.97% WER at peak versus 15.97% for full-audio fine-tuning.

In catastrophic forgetting evaluations on the source domain (DefinedAI base training set), the base model achieves 12.8% WER while standard ASR fine-tuning severely degrades source performance. In contrast, text-only and mixed-batch adaptation limit source-domain degradation while simultaneously improving target-domain accuracy from 8.0% down to 4.5%–6.4% WER.

System / ConditionBanking WER (%)Agriculture WER (%)Musical Instruments WER (%)
Base Model8.0241.0531.18
Text-Only Adaptation (τa=0\tau_a = 0)6.38--
Standard ASR Fine-Tuning (100% Audio)4.5516.0415.97
Mixed-Batch Adaptation (10% Target Audio)4.1915.9116.97

Limitations

The study is scoped to English-language domains and evaluates on a specific set of four specialized verticals (Banking, Insurance, Healthcare, Agriculture, Musical Instruments), leaving multilingual generalizability unverified. The compute requirements are constrained to a 3B-parameter LLM (Llama-3.2-3B-Instruct), and scaling behavior to larger 70B+ instruction models or alternative encoder architectures like Whisper remains unexplored. Additionally, the reliance on rule-based character-level perturbations via nlpaug for text corruption may not fully capture realistic acoustic-to-text projection errors in highly divergent vocabularies.

Why read this

Speech and ML engineers tackling low-resource domain adaptation for LLM-based ASR will find this paper a practical blueprint for exploiting abundant text data without suffering from modality mismatch. Readers will take away a validated batch composition recipe that eliminates the need to collect massive domain-specific audio corpora.

Code

Applications

Low-resource domain adaptation of speech recognition systems for specialized verticals such as banking, legal, and healthcare transcription where text documents are plentiful but audio recordings are rare.

Institutions

Idiap Research Institute, Laboratoire d’Informatique et des Systèmes, EPFL, Uniphore, Brno University of Technology

Funding / 經費: Idiap Research Institute, Uniphore, EU Horizon 2020

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3383