All papers
Speech translationFull-paper digest

POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation

Xuanchen Li, Chenrui Cui, Tianrui Wang, Meng Ge, Zikang Huang, Yizhou Peng, Jin Li, Yuheng Lu, Yu Jiang, Nyima Tashi, Longbiao Wang, Jianwu Dang

Code & resourcesgithub.com/Sslnon/POTSA

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.1 KB · Ready to paste

Preview copied content

TL;DR — POTSA is a cross-lingual alignment framework for SpeechLLMs that combines coarse bias compensation with token-level optimal transport on parallel speech pairs, achieving state-of-the-art speech-to-text translation using only 10 hours of parallel speech per language.

Key contributions

  • Proposes a lightweight Bias Compensation module that subtracts language-specific mean offsets from encoder outputs to establish a coarsely aligned shared representation space.
  • Applies entropy-regularized Optimal Transport (Sinkhorn distance) for token-level fine-grained cross-lingual alignment over parallel speech pairs without requiring strict point-to-point index matching.
  • Introduces an online reward-guided layer scheduling strategy using Upper Confidence Bound (UCB) and temperature-controlled sampling to restrict OT constraints to lower/beneficial Q-Former layers.
  • Demonstrates robust data efficiency, improving zero-shot translation performance by +2.93 BLEU and standard multilingual translation by +1.29 BLEU using only 10 hours of parallel data per language.

Problem

State-of-the-art Speech Large Language Models (SpeechLLMs) suffer from severe translation performance bias, exhibiting high accuracy on high-resource languages while lagging significantly behind on low-resource ones. Prior work predominantly relies on text-driven, unidirectional speech-to-text alignment, which compresses fine-grained acoustic details and treats each source language independently in language-specific clusters. This isolation prevents effective cross-lingual knowledge transfer and reuse of machine translation decoders, making explicit cross-lingual speech representation alignment essential.

Method

The architecture integrates within the SLAM-LLM framework using a frozen speech encoder (Whisper-v3) and a frozen language model (Qwen-2.5-7B), training only the Q-Former projection layer (8 Transformer blocks, 80 query tokens). The pipeline begins with a Bias Compensation module that models each encoder output as the sum of a language-neutral component and a language-specific bias vector computed via sentence-level temporal pooling and subtracted during both training and inference.

Following coarse alignment, the model applies token-level Optimal Transport (OT) using cross-lingual parallel speech pairs. For selected intermediate Q-Former layers, the Sinkhorn distance (entropy-regularized OT with squared Euclidean ground cost and transport smoothness parameter epsilon) is minimized to enforce semantic consistency between token sequences. This is jointly optimized with the standard translation cross-entropy loss using a weighting coefficient of 10.

To prevent alignment objectives from conflicting with translation supervision in deeper layers, an online reward-guided layer scheduling strategy based on Upper Confidence Bound (UCB) and temperature-controlled softmax sampling dynamically selects which lower Q-Former layers receive OT constraints based on reward moving averages.

Experimental setup

Pretrained on 364 hours of CoVoST2 data (ASR en->en followed by S2TT en->zh) and fine-tuned on FLEURS (~10 hours per language across 5 source languages: en, ja, es, ko, ru translating to zh). Evaluated against baselines including Qwen2-Audio, Qwen2.5-Omni, MinMo, and an end-to-end WhisperV3+Qwen2.5 baseline using BLEU scores, Recall@1, and 1-JSD on NVIDIA RTX 4090 GPUs.

Results

The proposed POTSA framework achieves an average BLEU score of 31.84 across the five training languages (+1.29 BLEU improvement over the baseline) and 20.84 on zero-shot languages (+2.93 BLEU improvement). Ablation studies confirm that combining both bias compensation and OT alignment is vital, as removing OT drops performance to 30.73 and removing bias compensation drops it to 31.11. Comparing alignment losses, OT outperforms MSE (30.52) and Cosine Similarity (30.74). Crucially, applying OT to all layers or deeper layers degrades performance due to objectives competing with the cross-entropy loss, validating the selective lower-layer scheduling strategy.

Systems / Conditionsen->zhja->zhes->zhko->zhru->zhAvg (5 Lang)Avg (Zero-Shot)
Qwen2-Audio [29]37.4520.0228.6925.6929.0128.1716.67
Qwen2.5-Omni [8]38.3523.7627.8127.0529.6929.3316.91
WhisperV3+Qwen2.5 (Baseline)40.0924.3230.1427.0731.1330.5517.91
POTSA (Our Model)40.8725.9731.1028.9732.3031.8420.84
w/o Bias Comp.40.4424.8330.9128.0631.3131.1119.34
w/o OT Align.40.2225.2230.8925.9131.4030.7318.51

Limitations

The framework relies on the availability of parallel speech pairs across source languages, which, while limited to 10 hours per language here, may still be scarce for extremely low-resource dialects. The evaluation is focused on translation into a single target language (Chinese), leaving many-to-many translation dynamics partially explored. Furthermore, freezing the speech encoder and LLM restricts adaptation to the parameters of the lightweight Q-Former.

Why read this

Speech and ML researchers tackling multilingual speech-to-text translation and modality alignment should read this to learn how to combine coarse bias subtraction with entropy-regularized optimal transport for robust cross-lingual representation learning.

Code

Applications

Real-world multilingual speech-to-text translation systems and cross-lingual spoken language understanding applications targeting low-resource languages.

Institutions

Tianjin University, Nanyang Technological University, Huiyan Technology Company, Tibet University, Chinese Academy of Sciences

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-695