All papers
Speech recognitionFull-paper digest

Exploring the potential and limitations of Model Merging for Multi-Domain Adaptation in ASR

Carlos Carvalho, Francisco Teixeira, Thomas Rolland, Alberto Abad

Code & resourcesgithub.com/Miamoto/mergewhisper

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.8 KB · Ready to paste

Preview copied content

TL;DR — The paper benchmarks 11 model merging algorithms for multi-domain European Portuguese ASR using Whisper Large-v3 and introduces BoostedTSV-M, achieving competitive in-domain accuracy while better preserving cross-lingual and out-of-distribution generalization compared to full joint fine-tuning.

Key contributions

  • Introduces MergeWhisper, an open-source extension of mergekit adding native support for Whisper architectures.
  • Systematically benchmarks 11 merging algorithms across 3 paradigms (parameter-space, tau-space, tau-subspace) on 10 European Portuguese domains.
  • Proposes BoostedTSV-M, mitigating rank collapse via singular-value boosting and substituting unstable orthogonal Procrustes with Newton-Schulz orthogonalization.
  • Demonstrates that model merging bypasses the catastrophic forgetting of multilingual capabilities typically seen in joint full fine-tuning.

Problem

Large speech foundation models are routinely adapted to new target domains via separate fine-tuning, yielding a proliferation of isolated checkpoints that complicate inference deployment. While joint fine-tuning on all data avoids checkpoint fragmentation, it requires access to historical training data that is frequently restricted by privacy or storage limits, and demands costly full retraining from scratch whenever new domains arrive. Sequential continual learning alternatives suffer from catastrophic forgetting and strong sensitivity to task presentation order. Model merging offers a non-sequential workaround by fusing independently trained domain checkpoints into a single parameter set, but its potential, trade-offs, and generalization capacity remain under-explored for speech foundation models.

Method

The paper evaluates parameter-space (SOUP, Model Stock, Karcher Mean, Multi-SLERP), tau-space (TA, TIES, PCB, SCE), and tau-subspace (ISO-C, ISO-CTS, TSV-M) methods applied on top of Whisper Large-v3. To overcome numerical instabilities when retaining high rank percentages (above 50%), the authors replace the standard orthogonal Procrustes algorithm in TSV-M and ISO-CTS with Newton-Schulz orthogonalization (using 5 iterations and a quintic coefficient schedule).

To combat rank collapse and the suppression of task-specific signals caused by sharply decaying singular values, the authors propose BoostedTSV-M. Prior to concatenating singular vectors and values across tasks, BoostedTSV-M calculates cumulative energy ratios and applies a lower-bound thresholding parameter beta (set to 0.3 after sweeping) to boost small singular values. The final merged model is constructed as theta_star = theta_0 + lambda * tau_BoostedTSV-M.

For training individual domain checkpoints (ID-FT), models are optimized for 3 epochs with a learning rate of 2e-6 and an effective batch size of 256. Full joint fine-tuning (Full-FT) runs for 10 epochs with a learning rate of 1e-5. All inferences use VAD-based segmentation via WhisperX with a beam size of 5 to mitigate hallucinations.

Experimental setup

Evaluated on 10 European Portuguese (EP) corpora totaling ~350 hours of training data and 46.2 hours of test data (including 10 OOD test splits, African/Asian Portuguese, Brazilian Portuguese, English via OpenASR-HF, and 21 languages of FLEURS). Systems compared include zero-shot Whisper Large-v3-X, Full-FT, ID-FT upper bounds, and 11 merging variants. Metrics are Word Error Rate (WER) and Character Error Rate (CER) for unsegmented languages, implemented via HuggingFace using NVIDIA A100 (80GB) and RTX A6000 (48GB) hardware.

Results

Full-FT drops in-domain EP WER from the zero-shot baseline of 15.62% down to 8.54% (and EP OOD from 25.21% to 17.65%), but causes degradation on English and multilingual FLEURS benchmarks due to catastrophic forgetting. BoostedTSV-M achieves the best overall EP average error rate of 11.55%, marginally beating Full-FT (11.58%) while substantially preserving out-of-distribution capabilities.

Parameter-space methods like Karcher Mean and Model Stock excel at non-EP OOD generalization (e.g., Model Stock achieves the best average FLEURS error rate of 6.79%), whereas tau-subspace methods dominate in-domain EP accuracy.

SystemEP ID WER (%)EP OOD WER (%)EP Full Avg. (%)OpenASR-HF WER (%)FLEURS Avg. (%)
Zero-shot15.6225.2118.827.177.43
Full-FT8.5417.6511.588.839.63
Karcher Mean10.2317.6312.707.127.43
TSV-M w/ NS9.4116.0711.637.249.68
BoostedTSV-M9.2716.1111.557.6010.37

Limitations

The study is restricted to European Portuguese and select language variants, leaving open whether the observed singular-value boosting hyperparameters translate effectively to highly tonyl or morphologically distant languages. The approach requires access to independently fine-tuned domain checkpoints and does not update or compress the base foundation model architecture itself. Furthermore, boosting task-specific singular values inherently trades away some cross-lingual and OOD transfer capabilities, exposing a friction point between specialization and generalization.

Why read this

Speech and ML engineers dealing with multi-domain ASR deployment constraints should read this to learn how to bypass costly joint re-training and avoid checkpoint sprawl without sacrificing out-of-distribution robustness.

Code

Applications

Multi-domain conversational agents, localized speech transcription systems, and enterprise speech recognition deployed across diverse regional dialects.

Institutions

INESC-ID, Universidade de Lisboa

Funding / 經費: Fundacao para a Ciencia e a Tecnologia, Portuguese Recovery and Resilience Plan

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1969