TL;DR — GigaAM Multilingual is a 600M-parameter Conformer ASR encoder pretrained on 2 million hours of audio, featuring cluster-level pretraining balancing and domain-aware fine-tuning sampling to drastically improve low-resource speech recognition. It outperforms much larger models like Whisper Large v3 and Omnilingual-1B on Central Asian languages (Kazakh, Kyrgyz, Uzbek).
Key contributions
- Introduces GigaAM Multilingual, a state-of-the-art open-weight Conformer ASR foundation model targeting underrepresented Central Asian languages (Kyrgyz, Kazakh, Uzbek).
- Proposes a cluster-level data balancing strategy during SSL pretraining based on a normalized language co-occurrence graph to mitigate head-language dominance.
- Develops a multi-domain fine-tuning recipe combining open-source data, crowdsourcing with ROVER quality control, synthetic speech, and weakly supervised forced alignment with domain-aware sampling.
- Demonstrates superior cross-lingual transfer and low-resource adaptation on minimal-coverage tail languages (Bashkir, Georgian) under controlled CTC fine-tuning compared to Whisper Large v3 and Omnilingual-1B.
Problem
Despite recent scaling advances in multilingual automatic speech recognition, performance remains severely uneven, with long-tail and low-resource languages exhibiting high error rates that block downstream deployment. This disparity is primarily driven by data imbalance, where naive upsampling of underrepresented languages either fails to overcome head-language dominance or degrades performance on high-resource languages. Prior foundational models like Whisper, Omnilingual, and Seamless M4T often struggle heavily in low-resource regimes such as Central Asian languages, highlighting the need for principled pretraining and adaptation data-mixing strategies.
Method
GigaAM Multilingual is built on a 600M-parameter Conformer encoder architecture consisting of 24 layers, a hidden dimension of 1024, and Rotary Position Embeddings in self-attention layers. The encoder operates at a frame rate of 25 Hz with a 40 ms stride, consuming input mel-spectrograms. Pretraining uses a HuBERT-style masked unit prediction objective, where a student encoder predicts discrete acoustic units at 40% randomly masked contiguous spans. These target units are generated by running K-means clustering (K = 1000) over teacher representations extracted from an in-house pretraining corpus of 2 million hours of audio spanning over 70 languages.
To manage data imbalance without manual per-language tuning, the authors construct a normalized language co-occurrence graph where edge weights reflect how frequently languages co-occur within the same audio recording. Pruning low-weight edges yields five distinct language clusters. Audio files are assigned to languages using voice activity detection, MMS LID 4017 language identification, and confidence thresholds (>= 0.7 majority vote and model confidence), allowing cluster-level sampling weights (Experiment E2 configuration: [0.50, 0.15, 0.25, 0.05, 0.05]) to upweight Central Asian groups (C3) during pretraining.
For downstream adaptation, the pretrained encoder is fine-tuned using a Connectionist Temporal Classification (CTC) objective over a shared character vocabulary spanning English, Russian, Kazakh, Kyrgyz, and Uzbek. The fine-tuning dataset combines 38k+ hours of open-source corpora, crowdsourced speech verified by 5-way annotator ROVER consensus, synthetic speech generated via an in-house multi-speaker TTS system (filtered via ASR verification with CER < threshold), and weakly supervised long-form recordings processed via MMS forced alignment. To optimize training, a domain-aware sampling strategy is employed to prevent massive synthetic subsets from dominating the curriculum and degrading spontaneous speech generalization.
Experimental setup
Pretraining uses 2 million hours of in-house audio, trained for 300k steps with a learning rate of 2e-4 and a total audio batch size of 2048 x 32 seconds per update. Fine-tuning runs for 200k steps using AdamW with a peak learning rate of 6e-5, a 5k-step warmup, cosine decay to 1e-7, and a virtual batch size of 3200 via gradient accumulation. Evaluation benchmarks include Common Voice, FLEURS, and internal in-the-wild crowdsourced datasets (6-20 hours per language), measured using Normalized Word Error Rate (WER). Baselines include Whisper Large v3, Omnilingual-1B LLM ASR, and Seamless M4T Large v2.
Results
Under controlled CTC fine-tuning on internal test sets, GigaAM Multilingual (600M) achieves dramatic WER reductions compared to open baselines: on Uzbek internal, GigaAM scores 12.7% WER vs. Omnilingual 1B's 30.2% and Seamless M4T's 40.0%; on Kazakh internal, GigaAM scores 15.8% vs. Omnilingual's 32.2% and Whisper's 65.2%; on Kyrgyz internal, GigaAM scores 9.8% vs. Omnilingual's 25.0% and Whisper's 102.2%. In pretraining cluster ablations, shifting weights toward the Central Asian cluster (E2) improved Kyrgyz WER from 9.4% to 8.5% and Uzbek from 10.5% to 9.7% with minimal degradation on high-resource languages.
In encoder ablations under matched CTC fine-tuning, even the compact 240M GigaAM variant outperformed Whisper Large v3 (average WER 12.2% vs 14.1%) and Omnilingual 1B (16.6%). On minimal-coverage tail languages with <500 hours of pretraining coverage (Georgian and Bashkir Common Voice), GigaAM achieved lowest WERs of 3.8% (Georgian) and 3.6% (Bashkir), substantially beating Omnilingual SSL (7.8% / 8.2%) and Whisper Large v3 CTC (13.4% / 11.1%).
| System | Russian (Int) | Kazakh (Int) | Kyrgyz (Int) | Uzbek (Int) |
|---|---|---|---|---|
| GigaAM Multilingual | 6.0 | 15.8 | 9.8 | 12.7 |
| Omnilingual 1B ASR | 14.6 | 32.2 | 25.0 | 30.2 |
| Seamless M4T large v2 | 16.1 | 62.9 | 78.3 | 40.0 |
| Whisper large v3 | 10.1 | 65.2 | 102.2 | 120.6 |
Limitations
The work focuses predominantly on three Central Asian languages (Kazakh, Kyrgyz, Uzbek) alongside Russian and English, leaving broader global low-resource language coverage largely untested. Synthetic data scaling relies heavily on pre-existing text corpora and quality-filtering thresholds that may introduce domain biases. Evaluation is restricted to greedy decoding on clean-to-spontaneous datasets up to 30 seconds long with digit-filtered references, limiting insights into long-form conversational robustness.
Why read this
Researchers and engineers tackling low-resource speech recognition will find a rigorously ablated blueprint for combining cluster-aware SSL pretraining weights with multi-domain fine-tuning data mixtures. It demonstrates how a mid-sized encoder (240M-600M parameters) can outperform massive foundation models when data imbalances are explicitly managed.
Code
Applications
Automated speech recognition and voice-driven applications for underrepresented Central Asian languages, including transcription of spontaneous speech and development of localized voice interfaces.
Institutions
SaluteDevices
Related
- Bootstrapping Endangered Language ASR with Short-Form Corpora — same problem · relatedness 2.3/3
- Genealogical Priors in Self-Supervised Learning: Improving Speech Technology for Low-Resource Languages — same problem · relatedness 2.3/3
- Alignment-Aware Continued Pre-training for Multilingual Speech Representation Learning — same problem · relatedness 2.2/3
- Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR — same problem · relatedness 2.1/3
- PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs — same problem · relatedness 2.1/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2483