All papers
Speaker recognitionFull-paper digest

BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

Théo Charlot, Tarek Kunze, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, Marvin Lavechin

Code & resourcesgithub.com/LAAC-LSCP/VTC

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.6 KB · Ready to paste

Preview copied content

TL;DR — BabyHuBERT is a multilingual self-supervised speech representation model trained on 13,164 hours of child-centered daylong recordings across 40+ languages, achieving a state-of-the-art average F1-score of 66.9% on Voice Type Classification (approaching human annotator performance of 69.8%).

Key contributions

  • Constructed the first massive-scale multilingual pre-training corpus for child-centered recordings spanning over 40 languages and 13,164 effective hours.
  • Introduced a robust speech preprocessing pipeline that extracts speech segments from daylong continuous audio, reducing non-speech content from ~80% to 8%.
  • Developed a two-iteration HuBERT pre-training recipe (leveraging WavLM features and custom transformer layer clustering) tailored for noisy, fragmented infant and child speech environments.
  • Achieved significant performance gains on Voice Type Classification, outperforming prior models by 13.3+ absolute F1 points and closing in on human baseline performance.

Problem

Automatic speech processing models trained predominantly on clean adult speech fail catastrophically on child-centered daylong recordings due to heavy acoustic noise, non-speech content (~80% of raw audio), fragmented vocalizations, overlapping speakers, and distinct acoustic properties of child speech. Prior domain-specific models like W2V2-LL4300 are limited by an English-only focus and smaller pre-training scale, hindering automated developmental research across diverse linguistic contexts.

Method

BabyHuBERT adopts the HuBERT base architecture (12 transformer layers) and a two-iteration masked prediction approach. Because raw daylong audio is ~80% silence or environmental noise, the authors used PyanNet-VTC to extract and merge speech segments, shortening chunks shorter than 2s with padding and capping chunks at 30s, reducing non-speech content to ~8%.

For the first iteration (BabyHuBERT-1), discrete pseudo-targets were generated by applying MiniBatchKMeans (500 clusters) on features extracted from the 6th layer of WavLM-base-plus. For the second iteration (BabyHuBERT-2), targets were generated by clustering features from the 7th transformer layer of BabyHuBERT-1. Both iterations were trained for 400k steps using torchaudio on 32 H100 GPUs with a batch size of 175 seconds per GPU (~85 effective seconds per GPU after bucketing), running for 45 and 44 epochs (~30 hours each).

For downstream Voice Type Classification (VTC), four independent linear classification heads with 0.5 dropout were added on top of the encoder's last layer to perform multi-label binary classification (Key Child, Other Children, Male Adult, Female Adult) to handle overlapping speech. Only the transformer layers were fine-tuned while convolutional feature extractors remained frozen, utilizing an initial learning rate of 1e-5 reduced on plateau with a batch size of 128 utterances (10 seconds each).

Experimental setup

Pre-training used 19 diverse child-centered corpora totaling 39,029 hours of raw audio (13,164 effective hours after filtering) covering over 40 languages. Fine-tuning and evaluation utilized the BabyTrain-2025 dataset (670 hours total: 158h KCHI, 11h OCH, 12h MAL, 262h FEM) split 80/10/10 child-disjoint, plus a 20-hour ACLEW hold-out set. Baselines include LENA, PyanNet-VTC, Whisper-VTC, HuBERT base/large (adult speech), and W2V2-LL4300 (English child speech). Models were evaluated using pyannote.metrics F1-score across 10 random seeds.

Results

BabyHuBERT-2 (BabyHuBERT-VTC) achieved an average F1-score of 66.9% on the hold-out set, outperforming Whisper-VTC (53.6%), PyanNet-VTC (50.9%), W2V2-LL4300 (58.4%), and standard HuBERT base (50.7%), coming within 2.9 points of a second human annotator (69.8%). Notably, it achieved a 56.1% F1-score on the highly challenging Other Children (OCH) class—a 25.6 absolute point improvement over previous state-of-the-art systems. In cross-corpora test evaluations, BabyHuBERT consistently outperformed W2V2-LL4300 and HuBERT across all six tested languages, displaying particularly strong performance on highly multilingual corpora like Vanuatu (76.1% F1) and the Solomon Islands (73.5% F1).

SystemKCHIOCHMALFEMAve.
LENA™54.928.537.242.640.8
PyanNet-VTC68.230.541.263.750.9
Whisper-VTC68.420.656.768.953.6
W2V2-LL430068.141.055.469.158.4
BabyHuBERT-VTC (ours)67.156.168.875.566.9
Second human annotator79.760.467.671.569.8

Limitations

The pre-training scale was constrained to HuBERT-base architecture due to high compute costs, precluding exhaustive exploration of larger model capacities and hyperparameter sweeps. Single-microphone field recordings limit ground-truth accuracy for distinguishing peer children due to acoustic overlap and reverberation. The models are released under a restrictive custom license prohibiting commercial use and surveillance to respect indigenous data sovereignty and privacy.

Why read this

Speech and ML researchers building systems for noisy, resource-constrained, or child-centered acoustic environments will find a definitive blueprint for large-scale multilingual self-supervised domain adaptation. It proves that scaling multilingual child-centered pre-training data unlocks massive downstream performance gains for speaker segmentation tasks.

Code

Applications

Automated developmental psychology research, voice type classification in child-centered home environments, peer interaction analysis, and downstream speech processing for underrepresented languages.

Institutions

École Normale Supérieure, École des Hautes Études en Sciences Sociales, CNRS, PSL University, Aix-Marseille University

Funding / 經費: Agence Nationale pour la Recherche, European Research Council, Simons Foundation International, Agence de l'Innovation de Défense

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2772