TL;DR — CoCoVSR is a cross-lingual compositional learning framework that adapts pretrained multilingual visual speech recognition models to code-switched scenarios using synthetic pseudo-code-switched clips and shared adapters, achieving state-of-the-art results on the CSLR benchmark.
Key contributions
- Proposes a cross-lingual compositional learning strategy (CoCo-MultiVSR) that automatically builds diverse pseudo code-switched training samples from existing monolingual video corpora without extra annotation or generative synthesis.
- Applies a parameter-efficient finetuning (PEFT) framework using lightweight adapters to equip multilingual VSR models with code-switching competence.
- Achieves state-of-the-art performance on the CSLR English-Chinese benchmark while minimizing performance degradation on monolingual seen (MultiVSR Chinese) and unseen (LRS2 English) datasets.
Problem
Code-switching is a pervasive feature of natural conversation, yet visual speech recognition (VSR) has largely ignored it because collecting large-scale, realistic code-switched visual data is extremely burdensome and expensive. Existing open benchmarks like CSLR contain only about 5K short English-Chinese utterances with limited lexical and switching diversity, meaning models often fail to generalize to unseen compositional configurations. While code-switched automatic speech recognition relies on heavy mixture-of-experts or language-specific adapters that introduce parameter overhead and require explicit routing, these approaches are poorly suited to the shared articulatory and viseme patterns common across languages in visual speech.
Method
The CoCoVSR framework builds upon a pretrained multilingual VSR backbone composed of a 3D-CNN and visual transformer blocks (VTP) for the front-end, followed by a Transformer encoder-decoder. To overcome data scarcity, the model trains on pseudo code-switched samples (CoCo-MultiVSR) generated by randomly sampling Mandarin and English monolingual video clips from datasets like MultiVSR and concatenating them in temporal sequences using uniform probability over order permutations. To retain original multilingual knowledge while learning code-switching, the method freezes backbone parameters and trains lightweight LoRA adapters (set at rank 32) applied before feed-forward and self-attention layers in the encoder, and after multi-head self-attention and feed-forward layers in the decoder.
Unlike ASR systems that use language-isolated modules, CoCoVSR employs a shared adapter architecture. This design decision leverages the strong cross-lingual commonality of articulatory movements (where single visemes map to multiple phonemes across different languages), avoiding complex routing while successfully connecting code-switched transcription with the pretrained backbone's multilingual capabilities. During training, the system optimizes cross-entropy loss over text token predictions, utilizing a batch size of 8 for 50 epochs with a peak learning rate of 1e-4 and a linear warmup over the first 10% of iterations. At inference, beam search is used with a beam width of 20.
Experimental setup
Experiments use the CSLR English-Chinese benchmark (~5K short utterances, 3 seconds each, resampled to 25 fps, 96x96 mouth crops) alongside a usable subset of MultiVSR (540K English and 26K Chinese clips) and the LRS2 English test set (1.2K samples) for domain transfer. Evaluation metrics include Character Error Rate (CER) for Chinese, Word Error Rate (WER) for English, and Mixture Error Rate (MER) for code-switched utterances. Implementation uses 4 NVIDIA RTX 6000 GPUs, utilizing S3FD for face detection and cropping.
Results
CoCoVSR achieves state-of-the-art performance on the CSLR benchmark, drastically lowering mixture error rate (MER) to 16.48% (compared to 37.54% for MoE+CTC and 256.46% for raw MultiVSR). On individual language components within CSLR, it reaches 16.69% Chinese CER and 16.23% English WER. In ablation studies, varying the CSLR-to-CoCo-MultiVSR data ratio shows that while training exclusively on CSLR gives slightly better pure CSLR metrics (12.93% CER), adding synthetic data is necessary to prevent catastrophic degradation on monolingual evaluation benchmarks like LRS2 (improving LRS2 WER from 98.97% down to 55.51% at a 1:1 ratio). Furthermore, using a shared adapter outperforms multiple language-specific adapters, yielding 16.60% CER vs. 19.73% CER on CSLR code-switching.
| System / Condition | CSLR CER (Zh) | CSLR WER (En) | CSLR MER | MultiVSR CER (Zh) | LRS2 WER (En) |
|---|---|---|---|---|---|
| MultiVSR [19] | 524.04 | 100.00 | 256.46 | 67.03 | 58.71 |
| MoE+CTC [15] | 33.01 | 45.37 | 37.54 | - | - |
| CoCoVSR (1:0 Data Ratio) | 12.93 | 10.53 | 12.00 | 90.12 | 98.97 |
| CoCoVSR (1:1 Data Ratio, Ours) | 16.69 | 16.23 | 16.48 | 84.21 | 55.51 |
| CoCoVSR (1:2 Data Ratio) | 19.44 | 18.59 | 19.12 | 82.97 | 53.58 |
Limitations
The approach relies heavily on the quality and domain coverage of existing monolingual datasets to construct synthetic code-switched examples, which may not capture natural conversational prosody shifts or co-articulation effects at real code-switching boundaries. Evaluation is currently constrained to a single language pair (English-Mandarin), leaving the generalization to other language combinations unverified. Additionally, performance on pure monolingual benchmarks like MultiVSR Chinese experiences some degradation compared to dedicated monolingual training when the synthetic data ratio is tuned for optimal code-switching.
Why read this
Speech and ML researchers working on multimodal tasks or low-resource visual speech recognition should read this paper to see how simple cross-lingual compositional data augmentation and shared parameter-efficient adapters can solve data scarcity in code-switching without complex routing modules.
Code
Applications
Multilingual video transcription, cross-lingual video subtitle generation, and inclusive human-computer interaction systems supporting natural code-switched speech.
Institutions
Ewha Womans University
Funding / 經費: National Research Foundation of Korea
Related
- Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR — same problem · relatedness 2.0/3
- Reinforcement Learning for Data-Efficient Code-Switched ASR — same problem · relatedness 1.9/3
- A Multimodal Semi-Supervised Framework for Automatic Construction of a Cross-Lingual Taigi Speech-Chinese Subtitle Corpus — shared technique · relatedness 1.9/3
- Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech — same problem · relatedness 1.9/3
- Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio — shared technique · relatedness 1.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1163