---
id: charlot26_interspeech
title: "BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting
  Speakers in Child-Centered Long-Form Recordings"
authors:
  - Théo Charlot
  - Tarek Kunze
  - Maxime Poli
  - Alejandrina Cristia
  - Emmanuel Dupoux
  - Marvin Lavechin
year: 2026
doi: 10.21437/Interspeech.2026-2772
isca_url: https://www.isca-archive.org/interspeech_2026/charlot26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/charlot26_interspeech.pdf
session: "CHILDSPACE: Child Home Interaction & Language Dynamics: Speech,
  Psychology, Affect, Computation, and Environments"
topics:
  - self-supervised
  - speaker-diarization
  - multilingual
category: speaker
labels:
  - multilingual
  - self-supervised
institutions:
  - École Normale Supérieure
  - École des Hautes Études en Sciences Sociales
  - CNRS
  - PSL University
  - Aix-Marseille University
funding:
  - Agence Nationale pour la Recherche
  - European Research Council
  - Simons Foundation International
  - Agence de l'Innovation de Défense
code:
  url: https://github.com/LAAC-LSCP/VTC
  stars: 15
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: charlot26_interspeech
  category: speaker
  labels:
    - multilingual
    - self-supervised
  institutions:
    - École Normale Supérieure
    - École des Hautes Études en Sciences Sociales
    - CNRS
    - PSL University
    - Aix-Marseille University
  code: https://github.com/LAAC-LSCP/VTC
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-2772
  pdf: https://www.isca-archive.org/interspeech_2026/charlot26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/charlot26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/charlot26_interspeech/markdown.md
---

# BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

*Théo Charlot, Tarek Kunze, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, Marvin Lavechin*

[PDF](https://www.isca-archive.org/interspeech_2026/charlot26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/charlot26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-2772)

**Category:** `speaker` · **Labels:** `multilingual`, `self-supervised`

**TL;DR** — BabyHuBERT is a multilingual self-supervised speech representation model trained on 13,164 hours of child-centered daylong recordings across 40+ languages, achieving a state-of-the-art average F1-score of 66.9% on Voice Type Classification (approaching human annotator performance of 69.8%).

## Key contributions

- Constructed the first massive-scale multilingual pre-training corpus for child-centered recordings spanning over 40 languages and 13,164 effective hours.
- Introduced a robust speech preprocessing pipeline that extracts speech segments from daylong continuous audio, reducing non-speech content from ~80% to 8%.
- Developed a two-iteration HuBERT pre-training recipe (leveraging WavLM features and custom transformer layer clustering) tailored for noisy, fragmented infant and child speech environments.
- Achieved significant performance gains on Voice Type Classification, outperforming prior models by 13.3+ absolute F1 points and closing in on human baseline performance.

## Problem

Automatic speech processing models trained predominantly on clean adult speech fail catastrophically on child-centered daylong recordings due to heavy acoustic noise, non-speech content (~80% of raw audio), fragmented vocalizations, overlapping speakers, and distinct acoustic properties of child speech. Prior domain-specific models like W2V2-LL4300 are limited by an English-only focus and smaller pre-training scale, hindering automated developmental research across diverse linguistic contexts.

## Method

BabyHuBERT adopts the HuBERT base architecture (12 transformer layers) and a two-iteration masked prediction approach. Because raw daylong audio is ~80% silence or environmental noise, the authors used PyanNet-VTC to extract and merge speech segments, shortening chunks shorter than 2s with padding and capping chunks at 30s, reducing non-speech content to ~8%. 

For the first iteration (BabyHuBERT-1), discrete pseudo-targets were generated by applying MiniBatchKMeans (500 clusters) on features extracted from the 6th layer of WavLM-base-plus. For the second iteration (BabyHuBERT-2), targets were generated by clustering features from the 7th transformer layer of BabyHuBERT-1. Both iterations were trained for 400k steps using torchaudio on 32 H100 GPUs with a batch size of 175 seconds per GPU (~85 effective seconds per GPU after bucketing), running for 45 and 44 epochs (~30 hours each).

For downstream Voice Type Classification (VTC), four independent linear classification heads with 0.5 dropout were added on top of the encoder's last layer to perform multi-label binary classification (Key Child, Other Children, Male Adult, Female Adult) to handle overlapping speech. Only the transformer layers were fine-tuned while convolutional feature extractors remained frozen, utilizing an initial learning rate of 1e-5 reduced on plateau with a batch size of 128 utterances (10 seconds each).

## Experimental setup

Pre-training used 19 diverse child-centered corpora totaling 39,029 hours of raw audio (13,164 effective hours after filtering) covering over 40 languages. Fine-tuning and evaluation utilized the BabyTrain-2025 dataset (670 hours total: 158h KCHI, 11h OCH, 12h MAL, 262h FEM) split 80/10/10 child-disjoint, plus a 20-hour ACLEW hold-out set. Baselines include LENA, PyanNet-VTC, Whisper-VTC, HuBERT base/large (adult speech), and W2V2-LL4300 (English child speech). Models were evaluated using pyannote.metrics F1-score across 10 random seeds.

## Results

BabyHuBERT-2 (BabyHuBERT-VTC) achieved an average F1-score of 66.9% on the hold-out set, outperforming Whisper-VTC (53.6%), PyanNet-VTC (50.9%), W2V2-LL4300 (58.4%), and standard HuBERT base (50.7%), coming within 2.9 points of a second human annotator (69.8%). Notably, it achieved a 56.1% F1-score on the highly challenging Other Children (OCH) class—a 25.6 absolute point improvement over previous state-of-the-art systems. In cross-corpora test evaluations, BabyHuBERT consistently outperformed W2V2-LL4300 and HuBERT across all six tested languages, displaying particularly strong performance on highly multilingual corpora like Vanuatu (76.1% F1) and the Solomon Islands (73.5% F1).

| System | KCHI | OCH | MAL | FEM | Ave. |
|---|---|---|---|---|---|
| LENA™ | 54.9 | 28.5 | 37.2 | 42.6 | 40.8 |
| PyanNet-VTC | 68.2 | 30.5 | 41.2 | 63.7 | 50.9 |
| Whisper-VTC | 68.4 | 20.6 | 56.7 | 68.9 | 53.6 |
| W2V2-LL4300 | 68.1 | 41.0 | 55.4 | 69.1 | 58.4 |
| BabyHuBERT-VTC (ours) | 67.1 | 56.1 | 68.8 | 75.5 | 66.9 |
| Second human annotator | 79.7 | 60.4 | 67.6 | 71.5 | 69.8 |

## Limitations

The pre-training scale was constrained to HuBERT-base architecture due to high compute costs, precluding exhaustive exploration of larger model capacities and hyperparameter sweeps. Single-microphone field recordings limit ground-truth accuracy for distinguishing peer children due to acoustic overlap and reverberation. The models are released under a restrictive custom license prohibiting commercial use and surveillance to respect indigenous data sovereignty and privacy.

## Why read this

Speech and ML researchers building systems for noisy, resource-constrained, or child-centered acoustic environments will find a definitive blueprint for large-scale multilingual self-supervised domain adaptation. It proves that scaling multilingual child-centered pre-training data unlocks massive downstream performance gains for speaker segmentation tasks.

## Code

- https://huggingface.co/MarvinLvn/BabyHuBERT

## Applications

Automated developmental psychology research, voice type classification in child-centered home environments, peer interaction analysis, and downstream speech processing for underrepresented languages.

## Institutions / 機構

École Normale Supérieure, École des Hautes Études en Sciences Sociales, CNRS, PSL University, Aix-Marseille University

**Funding / 經費:** Agence Nationale pour la Recherche, European Research Council, Simons Foundation International, Agence de l'Innovation de Défense

## Related

- [Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning](fan26b_interspeech.md) — same problem · relatedness 2.7/3
- [BiMamba2 Masked Discrete-Unit Prediction for Multilingual Speech Representation for Unsupervised Speech in the Wild Challenge](subedi26_interspeech.md) — same problem · relatedness 2.2/3
- [Multi-Speaker Embeddings With Weakly Supervised Speaker Activity Detection For Granular Speaker Diarization](thienpondt26_interspeech.md) — same problem · relatedness 2.2/3
- [Speaker Separation via Audio Language Modeling](lanzendoerfer26b_interspeech.md) — same problem · relatedness 2.1/3
- [SDR-LLM: Speech-LLM Based End-to-End Speaker Diarization and Recognition with Sentence-Level Temporal Modeling](yu26g_interspeech.md) — same problem · relatedness 2.1/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
