---
id: banerasroux26_interspeech
title: Closing the Speech-Text Gap with Limited Audio for Effective Domain
  Adaptation in LLM-Based ASR
authors:
  - Thibault Bañeras-Roux
  - Sergio Burdisso
  - Esaú Villatoro-Tello
  - Dairazalia Sánchez-Cortés
  - Shiran Liu
  - Severin Baroudi
  - Shashi Kumar
  - Hasindri Watawana
  - Manjunath K E
  - Kadri Hacioglu
  - Petr Motlicek
  - Andreas Stolcke
year: 2026
doi: 10.21437/Interspeech.2026-3383
isca_url: https://www.isca-archive.org/interspeech_2026/banerasroux26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/banerasroux26_interspeech.pdf
session: Multimodal Speech Processing and Speech LLM Systems
topics:
  - asr
  - self-supervised
  - speech-llm
category: asr
labels:
  - low-resource
institutions:
  - Idiap Research Institute
  - Laboratoire d’Informatique et des Systèmes
  - EPFL
  - Uniphore
  - Brno University of Technology
funding:
  - Idiap Research Institute
  - Uniphore
  - EU Horizon 2020
code:
  url: https://github.com/idiap/llm-asr-text-only-adaptation
  stars: 5
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: banerasroux26_interspeech
  category: asr
  labels:
    - low-resource
  institutions:
    - Idiap Research Institute
    - Laboratoire d’Informatique et des Systèmes
    - EPFL
    - Uniphore
    - Brno University of Technology
  code: https://github.com/idiap/llm-asr-text-only-adaptation
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-3383
  pdf: https://www.isca-archive.org/interspeech_2026/banerasroux26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/banerasroux26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/banerasroux26_interspeech/markdown.md
---

# Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

*Thibault Bañeras-Roux, Sergio Burdisso, Esaú Villatoro-Tello, Dairazalia Sánchez-Cortés, Shiran Liu, Severin Baroudi, Shashi Kumar, Hasindri Watawana, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke*

[PDF](https://www.isca-archive.org/interspeech_2026/banerasroux26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/banerasroux26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-3383)

**Category:** `asr` · **Labels:** `low-resource`

**TL;DR** — This paper investigates domain adaptation for LLM-based ASR when target-domain audio is scarce but text is abundant, proposing a mixed batching strategy that combines text-only data with a small fraction of speech. Using only 10% target-domain speech (under 4 hours) in mixed batches achieves word error rates comparable to or better than conventional fine-tuning with 100% of the speech data.

## Key contributions

- Analyzed the trade-off between target-domain specialization and source-domain retention as a function of target data proportions in LLM-based ASR batch composition.
- Proposed a hybrid mixed-batch adaptation method integrating limited paired speech-text data into a text-heavy adaptation loop to bridge the cross-modal alignment gap.
- Demonstrated that the mixed batching strategy reduces catastrophic forgetting on source domains while boosting target-domain recognition compared to standard ASR fine-tuning.
- Released open-source code for reproducibility of the proposed domain adaptation pipelines.

## Problem

Conventional end-to-end and LLM-based automatic speech recognition systems rely heavily on large amounts of paired speech-text data for domain adaptation, which is expensive and scarce in specialized verticals like healthcare, finance, or agriculture. While adapting LLMs using abundant in-domain text-only data is a scalable alternative, it introduces a severe modality gap because the LLM unlearns the noisy acoustic representations generated by upstream speech encoders. Prior text-only adaptation methods rely on soft prompting or decoding tweaks, failing to evaluate whether minimal amounts of target-domain audio can preserve cross-modal alignment without requiring full speech collection.

## Method

The architecture follows the SLAM-ASR framework, comprising a frozen WavLM-Large speech encoder, a linear speech projector downsampling acoustic frames by factor k, and a Llama-3.2-3B-Instruct LLM. The baseline model is first pretrained on source-domain paired speech-text data from DefinedAI (Banking, Insurance, and Healthcare partitions, totaling 38h 10m) where only the speech projector is fine-tuned while the encoder remains frozen. Domain adaptation is then performed by updating only LoRA adapters inserted into the LLM for 5 epochs with a batch size of 10 under a fixed instruction prompt template.

The mixed batching (MB) strategy constructs training batches by dividing data into source and target components. Source data comprises three segments: (1) paired source speech-text ($\sigma_a$), (2) projector-induced noisy tokens mapped from speech projections to nearest LLM embeddings ($\sigma_{ta}$), and (3) synthetically corrupted source transcripts generated via random character-level perturbations using nlpaug ($\sigma_t$). The target portion consists of: (1) paired target speech-text ($\tau_a$), and (2) synthetically corrupted target transcripts ($\tau_t$) using identical character perturbations. The total target proportion $\tau = \tau_t + \tau_a$ is set to 50% based on sweep experiments showing optimal performance at this threshold, balancing underfitting/catastrophic forgetting against overfitting.

## Experimental setup

Experiments use the DefinedAI corpus for in-domain evaluation (Banking partition: 36h 20m train, 2h 08m dev, 4h 16m test) and SlideSpeech for out-of-distribution evaluation (Agriculture: 29h 20m train; Musical Instruments: 8h 30m train). Evaluations compare base models, text-only adaptation, standard ASR fine-tuning (paired speech-text only), and mixed-batch adaptation across varying fractions of target audio. The primary evaluation metric is Word Error Rate (WER). Models use Llama-3.2-3B-Instruct and WavLM-Large, trained for 5 epochs with batch size 10.

## Results

On the DefinedAI Banking in-domain test set, standard ASR adaptation yields a WER of 4.55% when using 100% of the speech data, whereas text-only adaptation reaches 6.38%, demonstrating a 1.83% absolute WER penalty (approx. 29% relative degradation) due to the modality gap. However, mixed-batch adaptation achieves competitive or superior performance using only 10% of the target speech data (approx. 3.5 hours for Banking, yielding WERs around 4.19%). For out-of-domain evaluation on SlideSpeech Agriculture, mixed-batch adaptation with 10% target audio matches or beats the 100% audio baseline (15.91% vs 16.04% WER). On Musical Instruments, mixed batching yields 16.97% WER at peak versus 15.97% for full-audio fine-tuning.

In catastrophic forgetting evaluations on the source domain (DefinedAI base training set), the base model achieves 12.8% WER while standard ASR fine-tuning severely degrades source performance. In contrast, text-only and mixed-batch adaptation limit source-domain degradation while simultaneously improving target-domain accuracy from 8.0% down to 4.5%–6.4% WER.

| System / Condition | Banking WER (%) | Agriculture WER (%) | Musical Instruments WER (%) |
|---|---|---|---|
| Base Model | 8.02 | 41.05 | 31.18 |
| Text-Only Adaptation ($\tau_a = 0$) | 6.38 | - | - |
| Standard ASR Fine-Tuning (100% Audio) | 4.55 | 16.04 | 15.97 |
| Mixed-Batch Adaptation (10% Target Audio) | 4.19 | 15.91 | 16.97 |

## Limitations

The study is scoped to English-language domains and evaluates on a specific set of four specialized verticals (Banking, Insurance, Healthcare, Agriculture, Musical Instruments), leaving multilingual generalizability unverified. The compute requirements are constrained to a 3B-parameter LLM (Llama-3.2-3B-Instruct), and scaling behavior to larger 70B+ instruction models or alternative encoder architectures like Whisper remains unexplored. Additionally, the reliance on rule-based character-level perturbations via nlpaug for text corruption may not fully capture realistic acoustic-to-text projection errors in highly divergent vocabularies.

## Why read this

Speech and ML engineers tackling low-resource domain adaptation for LLM-based ASR will find this paper a practical blueprint for exploiting abundant text data without suffering from modality mismatch. Readers will take away a validated batch composition recipe that eliminates the need to collect massive domain-specific audio corpora.

## Code

- https://github.com/idiap/llm-asr-text-only-adaptation

## Applications

Low-resource domain adaptation of speech recognition systems for specialized verticals such as banking, legal, and healthcare transcription where text documents are plentiful but audio recordings are rare.

## Institutions / 機構

Idiap Research Institute, Laboratoire d’Informatique et des Systèmes, EPFL, Uniphore, Brno University of Technology

**Funding / 經費:** Idiap Research Institute, Uniphore, EU Horizon 2020

## Related

- [Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR](magoshi26_interspeech.md) — same problem · relatedness 2.4/3
- [TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs](peng26b_interspeech.md) — same problem · relatedness 2.4/3
- [Avoiding Catastrophic Forgetting in Text-Only Adaptation of LLM-based ASR via Multi-View Text Denoising](burdisso26_interspeech.md) — same problem · relatedness 2.4/3
- [Merging the Knowledge of LLMs for Automatic Speech Recognition](futami26_interspeech.md) — same problem · relatedness 2.3/3
- [Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains](wang26o_interspeech.md) — same problem · relatedness 2.1/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
