All papers
Speech LLMs & dialogueFull-paper digest

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?

Tomoya Mizumoto, Yusuke Fujita

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.1 KB · Ready to paste

Preview copied content

TL;DR — This paper investigates how incorporating bidirectional speech translation into speech encoder pre-training bridges the representational gap between audio encoders and text-based LLMs, leading to superior downstream performance in Speech LLMs. The bidirectional translation pre-training configuration achieves substantial error reductions in ASR and boosts translation BLEU scores across both seen and unseen languages.

Key contributions

  • Identifies that standard unidirectional translation pre-training (X to English, like Whisper) leaves English representations tied to surface phonetic forms because English audio is restricted to monolingual transcription.
  • Proposes a symmetric, bidirectional English translation pre-training objective (X to/from English) combined with a redesigned multi-target decoder prompt format.
  • Demonstrates that bidirectional translation pre-training consistently improves downstream ASR, speech translation, and intent classification across 1B and 3B Llama 3.2 based Speech LLMs.
  • Proves that the advantages of bidirectional pre-training persist even when the speech encoder is un-frozen and jointly fine-tuned during downstream Speech LLM training.

Problem

Connecting pre-trained speech encoders to text-based LLMs via a lightweight adaptor causes a structural misalignment because SSL and ASR-based encoders produce language-specific representations organized by acoustic and phonetic similarities, whereas LLMs operate in a language-agnostic semantic space. Prior models like Whisper use unidirectional translation (X to English) and train English exclusively on transcription, meaning English audio fails to leverage semantic abstraction. This forces lightweight adaptors to bridge a massive modality and structural gap, degrading cross-modal integration.

Method

The speech encoder architecture adopts the Whisper medium encoder. To handle both transcription and bidirectional translation, a multi-target decoder prompt format is introduced where the target language token precedes the task token (acting as a pre-condition), and the source language token is positioned after the task token (acting as a predicted property, e.g., <|BOS|><|en|><|translate|><|de|>). The pre-training corpora comprise 130k hours across English (73.6k h), Japanese (36.2k h), German (10.0k h), and Chinese (9.8k h), augmented with synthetic parallel text generated by Qwen2.5-32B-Instruct. Three task mixtures are evaluated: ASR-only (100% transcription), ASR & ST (X to en; non-English uses a 75:25 transcription-to-translation ratio, English is 100% transcription), and ASR & ST (X bidirectional en; 75:25 ratio applied uniformly to all languages).

The Speech LLM architecture links the pre-trained speech encoder to a frozen Llama-3.2-1B-Instruct or 3B-Instruct model using a lightweight trainable adaptor (a two-layer CNN for temporal downsampling followed by a linear projection). Base model pre-training runs for 3 epochs with a global batch size of 512 using 16 NVIDIA H100 GPUs, utilizing a piecewise-linear warmup (peak LR 2e-4) followed by cosine decay. The Speech LLM adaptor is trained for 25k steps (batch size 512, peak LR 1e-4) on 8 NVIDIA H100 GPUs using a multi-task dataset totaling 6.2k hours spanning ASR, ST, intent classification, and emotion recognition.

Experimental setup

Evaluated on 130k hours of pre-training audio (LibriSpeech, ReazonSpeech, Multilingual LibriSpeech, WenetSpeech, YODAS, Common Voice) across English, Japanese, German, and Chinese, and a 6.2k-hour downstream multi-task dataset (VoxPopuli, FLEURS, AISHELL, JSUT, CoVoST2, SpeechBSD, SLURP, Speech-MASSIVE, MELD). Compared against ASR-Only and ASR & ST (X to en) baselines. Evaluated on WER/CER for ASR, BLEU for Speech Translation (seen pairs and unseen target languages fa, id, sv, tr), and Accuracy for Intent Classification (SLURP, Speech-MASSIVE) and Emotion Recognition (MELD) using 1B and 3B Llama 3.2 LLM backbones.

Results

For the 3B LLM, ASR & ST (X bidirectional en) reduces German ASR CER from 42.5 (ASR-Only) down to 24.3, and improves English ASR WER from 11.6 to 11.0. For X to en translation on FLEURS, the bidirectional model achieves higher average BLEU scores, and for en to X translation, it reaches up to 30.9 BLEU on German FLEURS compared to 28.7 for ASR-Only. On downstream intent classification with the 3B model, English SLURP/Speech-MASSIVE accuracy jumps from 57.3 (ASR-Only) to 64.5 with bidirectional pre-training. Emotion recognition (MELD) remains largely unaffected (~49.2 to 49.5), showing that acoustic-sensitive tasks are not degraded by semantic abstraction. When un-freezing the encoder, the bidirectional configuration maintains its performance superiority (e.g., 22.9 average X to en BLEU vs 20.1 for ASR-Only).

Pre-training Task Mixtureen ASR (WER)ja ASR (CER)zh ASR (CER)de ASR (CER)en Intent (Acc)de Intent (Acc)
ASR Only11.622.424.342.557.357.9
ASR & ST (X -> en)11.616.621.426.258.562.1
ASR & ST (X <-> en)11.015.821.124.364.566.3

Limitations

Evaluated on a restricted subset of four primary languages (English, Japanese, Chinese, German) during pre-training due to compute constraints, omitting dozens of other languages covered by larger foundation models. The translation training data relies heavily on synthetic text generated by a 32B LLM (Qwen2.5-32B-Instruct) rather than purely human-annotated parallel speech corpora. Performance on acoustic-bound tasks like emotion recognition shows no direct benefit from semantic translation objectives, indicating limitations when tasks require fine-grained paralinguistic rather than abstract semantic features.

Why read this

Researchers and engineers building speech LLMs who want to understand how encoder pre-training objectives affect cross-modal alignment should read this paper to see how adding bidirectional translation bypasses the limitations of unidirectional Whisper-style models.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Multilingual Speech LLMs, spoken dialogue agents, real-time speech translation systems, and cross-lingual spoken language understanding applications.

Institutions

SB Intuitions

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3241