All papers
Resources & evaluationFull-paper digest

Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech

Kirill Borodin, Nikita Vasiliev, Vasiliy Kudryavtsev, Maxim Maslov, Mikhail Gorodnichev, Grach Mkrtchian

Code & resourcesgithub.com/lab260ru/balalaika

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

14.0 KB · Ready to paste

Preview copied content

TL;DR — Balalaika is a data-centric, modular open-source pipeline for processing raw Russian audio into a 5,100-hour prosody-aware corpus, yielding consistent performance gains in speech denoising and TTS under equalized training budgets.

Key contributions

  • A modular, open-source pipeline combining semantic VAD, multi-ASR ROVER consensus, and multi-stage quality/speaker-purity filtering.
  • Linguistic and prosodic text enrichment incorporating punctuation restoration, context-aware lexical stress assignment, e/yo normalization, and IPA phonemization via a custom G2P model.
  • A 5,078-hour multi-source Russian speech dataset (Balalaika) accompanied by complete annotation provenance and replication scripts.
  • Comprehensive empirical validation demonstrating that models trained on Balalaika data outperform those trained on 11 public Russian corpora under equalized training budgets for both speech denoising and TTS tasks.

Problem

Publicly available speech datasets for Russian suffer from simplistic annotation, reliance on monotonous audiobooks that lack conversational prosody, or poor handling of complex linguistic phenomena such as vowel reduction, palatalization, and mobile stress. Prevailing pipelines optimized for high-resource languages ignore these morphological and prosodic nuances, resulting in unnatural text-to-speech synthesis and poor speech processing model performance. Furthermore, web audio mining is hindered by a lack of automated, scalable annotation frameworks that can simultaneously handle semantic segmentation, multi-speaker filtering, and high-fidelity transcription for under-resourced languages.

Method

The Balalaika framework processes raw audio sequentially through eight specialized modules. First, audio is segmented using SmartTurnV3.1 semantic VAD to preserve context, discarding segments where speech share is below 70% or internal silence exceeds 1 second, and grouping consecutive chunks into 5 to 15-second windows. Second, automatic quality and speaker-purity screening removes clips shorter than 3 seconds, high impulsive-energy items with a CREST-factor > 10, low-perceptual-quality samples using NISQA-S MOS < 4.2, and multi-speaker or overlapping speech via pyannote diarization. Third, transcription generates five distinct hypotheses per segment (GigaAM-CTC-v3 plain, GigaAM-CTC-v3 with an n-gram language model for lexical priors and timestamps, GigaAM-RNNT-v3, Vosk, and T-one) which are fused using ROVER consensus decoding to minimize word errors. Fourth, timestamp extraction leverages the CTC+LM hypothesis to assign word-level boundaries.

Fifth, text streams are enriched for prosody by restoring sentence-final and intra-sentential punctuation using RuPunctBig as a proxy for phrase breaks. Sixth, context-aware lexical stress and e/yo normalization are applied using RuAccent to resolve homographs and pronunciation ambiguities. Seventh, text is converted to IPA phoneme sequences using a lightweight transformer encoder-decoder G2P model (d_model=128, d_ff=512, 3 encoder/decoder layers, 4 attention heads) trained on Wiktextract-derived IPA inventories using AdamW (lr 3e-4, batch size 256, label smoothing 0.1 for 10 epochs). Eighth, the compiled multi-layer annotated corpus integrates multi-source public Russian repositories into a unified 5,078-hour resource.

Experimental setup

Experiments evaluate a 25-hour subset of Balalaika against 11 public Russian corpora including DeepSpeech, GOLOS-C/F, M-AILABS, OpenSTT, RuLS, RUSLAN, Common Voice (MCV), and SOVA variants. Speech denoising experiments train SEMamba from scratch using an identical budget (Adam, lr=5e-4, batch size 8, 50k steps) with MUSAN noise and RIR augmentation, evaluated on a 3,000-sample test benchmark using CSIG, CBAK, COVL, PESQ, VISQOL, STOI, and SI-SDR. TTS experiments train a VITS model from scratch under an equalized budget (Adam, lr=1e-4, batch size 32, 100k steps), evaluated on a held-out set of 2,000 texts using NISQA, UTMOS, Character Error Rate (CER), and human evaluation (MOS and IntMOS via LabelSpeech with 7 raters per clip).

Results

Balalaika achieves top performance across objective perceptual metrics and human evaluations. For RQ1, the unified dataset outperforms all 11 baseline corpora, scoring highest in NISQA MOS (NMOS 4.484 vs M-AILABS 3.530 and RuLS 3.788), UTMOS (3.019 vs RuLS 2.816), and human MOS (4.601). For RQ2, SEMamba trained on Balalaika achieves top scores in CSIG (3.856), CBAK (3.165), COVL (3.340), PESQ (2.723), and SI-SDR (8.809 dB) compared to all baseline-trained denoisers. For RQ3, VITS trained on Balalaika achieves the highest TTS MOS (3.570), UTMOS (2.738), and human MOS (3.618), while maintaining a competitive CER of 0.1062 (second only to single-speaker RUSLAN's 0.0496, which leads in IntMOS at 3.182 vs Balalaika's 2.532). Ablations confirm that combining stress and punctuation yields the lowest CER and highest naturalness, and that raising the quality filter threshold from MOS > 3.5 to MOS > 4.2 improves both intelligibility and prosodic naturalness.

DatasetNMOSTTS MOSUTMOSMOS ± 95% CICER
M-AILABS [10]3.5303.0322.3072.962 ± 0.0520.0908
RUSLAN [37]3.7441.9132.1303.253 ± 0.0680.0496
RuLS [11]3.7882.9272.1072.750 ± 0.0940.1003
MCV [38]3.7563.2102.1232.749 ± 0.0620.2380
SOVA AB [12]2.9692.5571.4901.354 ± 0.0630.9112
Balalaika (ours)4.4843.5702.7383.618 ± 0.0830.1062

Limitations

Models were trained under fixed data and compute budgets without running to full convergence, meaning some baseline models may be undertrained. The pipeline relies heavily on language-dependent components (ASR models, punctuation restorers, stress placers, and G2P tools) specific to Russian, limiting direct zero-shot transfer to other languages without tool replacement. Furthermore, because several source corpora are represented in Balalaika's multi-source mixture, the test evaluation set shares partial domain overlap with training data, preventing a fully source-independent evaluation.

Why read this

Researchers and engineers building Russian speech generation systems or looking to construct automated, multi-layer web audio annotation pipelines should read this paper to adopt its robust ROVER fusion, prosody-enrichment, and data-filtering recipes.

Code

Applications

Training high-naturalness text-to-speech (TTS) systems, robust speech denoising models, and automated large-scale speech corpus mining pipelines for morphologically complex languages.

Institutions

Moscow Technical University of Communications and Informatics, BitmanagerAI

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-83