All papers
Speech recognitionFull-paper digest

GigaAM Multilingual: Foundation Model for Underrepresented Languages

Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov, Georgii Gospodinov, Dmitrii Bolotov, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin

Code & resourcesgithub.com/salute-developers/GigaAM

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

14.1 KB · Ready to paste

Preview copied content

TL;DR — GigaAM Multilingual is a 600M-parameter Conformer ASR encoder pretrained on 2 million hours of audio, featuring cluster-level pretraining balancing and domain-aware fine-tuning sampling to drastically improve low-resource speech recognition. It outperforms much larger models like Whisper Large v3 and Omnilingual-1B on Central Asian languages (Kazakh, Kyrgyz, Uzbek).

Key contributions

  • Introduces GigaAM Multilingual, a state-of-the-art open-weight Conformer ASR foundation model targeting underrepresented Central Asian languages (Kyrgyz, Kazakh, Uzbek).
  • Proposes a cluster-level data balancing strategy during SSL pretraining based on a normalized language co-occurrence graph to mitigate head-language dominance.
  • Develops a multi-domain fine-tuning recipe combining open-source data, crowdsourcing with ROVER quality control, synthetic speech, and weakly supervised forced alignment with domain-aware sampling.
  • Demonstrates superior cross-lingual transfer and low-resource adaptation on minimal-coverage tail languages (Bashkir, Georgian) under controlled CTC fine-tuning compared to Whisper Large v3 and Omnilingual-1B.

Problem

Despite recent scaling advances in multilingual automatic speech recognition, performance remains severely uneven, with long-tail and low-resource languages exhibiting high error rates that block downstream deployment. This disparity is primarily driven by data imbalance, where naive upsampling of underrepresented languages either fails to overcome head-language dominance or degrades performance on high-resource languages. Prior foundational models like Whisper, Omnilingual, and Seamless M4T often struggle heavily in low-resource regimes such as Central Asian languages, highlighting the need for principled pretraining and adaptation data-mixing strategies.

Method

GigaAM Multilingual is built on a 600M-parameter Conformer encoder architecture consisting of 24 layers, a hidden dimension of 1024, and Rotary Position Embeddings in self-attention layers. The encoder operates at a frame rate of 25 Hz with a 40 ms stride, consuming input mel-spectrograms. Pretraining uses a HuBERT-style masked unit prediction objective, where a student encoder predicts discrete acoustic units at 40% randomly masked contiguous spans. These target units are generated by running K-means clustering (K = 1000) over teacher representations extracted from an in-house pretraining corpus of 2 million hours of audio spanning over 70 languages.

To manage data imbalance without manual per-language tuning, the authors construct a normalized language co-occurrence graph where edge weights reflect how frequently languages co-occur within the same audio recording. Pruning low-weight edges yields five distinct language clusters. Audio files are assigned to languages using voice activity detection, MMS LID 4017 language identification, and confidence thresholds (>= 0.7 majority vote and model confidence), allowing cluster-level sampling weights (Experiment E2 configuration: [0.50, 0.15, 0.25, 0.05, 0.05]) to upweight Central Asian groups (C3) during pretraining.

For downstream adaptation, the pretrained encoder is fine-tuned using a Connectionist Temporal Classification (CTC) objective over a shared character vocabulary spanning English, Russian, Kazakh, Kyrgyz, and Uzbek. The fine-tuning dataset combines 38k+ hours of open-source corpora, crowdsourced speech verified by 5-way annotator ROVER consensus, synthetic speech generated via an in-house multi-speaker TTS system (filtered via ASR verification with CER < threshold), and weakly supervised long-form recordings processed via MMS forced alignment. To optimize training, a domain-aware sampling strategy is employed to prevent massive synthetic subsets from dominating the curriculum and degrading spontaneous speech generalization.

Experimental setup

Pretraining uses 2 million hours of in-house audio, trained for 300k steps with a learning rate of 2e-4 and a total audio batch size of 2048 x 32 seconds per update. Fine-tuning runs for 200k steps using AdamW with a peak learning rate of 6e-5, a 5k-step warmup, cosine decay to 1e-7, and a virtual batch size of 3200 via gradient accumulation. Evaluation benchmarks include Common Voice, FLEURS, and internal in-the-wild crowdsourced datasets (6-20 hours per language), measured using Normalized Word Error Rate (WER). Baselines include Whisper Large v3, Omnilingual-1B LLM ASR, and Seamless M4T Large v2.

Results

Under controlled CTC fine-tuning on internal test sets, GigaAM Multilingual (600M) achieves dramatic WER reductions compared to open baselines: on Uzbek internal, GigaAM scores 12.7% WER vs. Omnilingual 1B's 30.2% and Seamless M4T's 40.0%; on Kazakh internal, GigaAM scores 15.8% vs. Omnilingual's 32.2% and Whisper's 65.2%; on Kyrgyz internal, GigaAM scores 9.8% vs. Omnilingual's 25.0% and Whisper's 102.2%. In pretraining cluster ablations, shifting weights toward the Central Asian cluster (E2) improved Kyrgyz WER from 9.4% to 8.5% and Uzbek from 10.5% to 9.7% with minimal degradation on high-resource languages.

In encoder ablations under matched CTC fine-tuning, even the compact 240M GigaAM variant outperformed Whisper Large v3 (average WER 12.2% vs 14.1%) and Omnilingual 1B (16.6%). On minimal-coverage tail languages with <500 hours of pretraining coverage (Georgian and Bashkir Common Voice), GigaAM achieved lowest WERs of 3.8% (Georgian) and 3.6% (Bashkir), substantially beating Omnilingual SSL (7.8% / 8.2%) and Whisper Large v3 CTC (13.4% / 11.1%).

SystemRussian (Int)Kazakh (Int)Kyrgyz (Int)Uzbek (Int)
GigaAM Multilingual6.015.89.812.7
Omnilingual 1B ASR14.632.225.030.2
Seamless M4T large v216.162.978.340.0
Whisper large v310.165.2102.2120.6

Limitations

The work focuses predominantly on three Central Asian languages (Kazakh, Kyrgyz, Uzbek) alongside Russian and English, leaving broader global low-resource language coverage largely untested. Synthetic data scaling relies heavily on pre-existing text corpora and quality-filtering thresholds that may introduce domain biases. Evaluation is restricted to greedy decoding on clean-to-spontaneous datasets up to 30 seconds long with digit-filtered references, limiting insights into long-form conversational robustness.

Why read this

Researchers and engineers tackling low-resource speech recognition will find a rigorously ablated blueprint for combining cluster-aware SSL pretraining weights with multi-domain fine-tuning data mixtures. It demonstrates how a mid-sized encoder (240M-600M parameters) can outperform massive foundation models when data imbalances are explicitly managed.

Code

Applications

Automated speech recognition and voice-driven applications for underrepresented Central Asian languages, including transcription of spontaneous speech and development of localized voice interfaces.

Institutions

SaluteDevices

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2483