All papers
Speaker recognitionFull-paper digest

Self-supervised Speaker Verification with High-Confidence Pseudo-Label Selection and DINO-Style Self-Distillation Based on Pre-trained Models

Yishuang Li, Yi Yu, Wanli Dong, Weihao Gan

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.4 KB · Ready to paste

Preview copied content

TL;DR — A self-supervised speaker verification method combining multi-layer pre-trained model clustering consensus for high-confidence pseudo-labeling with DINO-style teacher-student self-distillation for low-confidence data, achieving a 1.09% EER on VoxCeleb1-O.

Key contributions

  • High-confidence pseudo-label selection (HCPLS) using multi-layer clustering consistency and Hungarian matching across speaker-discriminative pre-trained Transformer layers.
  • A joint training framework coupling pseudo-supervised learning (with label noise correction) on high-confidence data and DINO-style self-distillation (with confidence gating) on low-confidence data.
  • Iterative refinement procedure leveraging chunk-level offline label correction, reaching an EER of 1.09% on VoxCeleb1-O in 4 iterations.

Problem

Current high-performance speaker verification relies on costly, large-scale labeled datasets, which are constrained by privacy laws. Existing self-supervised speaker verification methods suffer from pseudo-label noise caused by single-layer clustering imperfections or struggle with data inefficiency by discarding low-confidence or ambiguous samples. Prior contrastive methods also risk treating same-speaker negatives as false negatives, while naive pseudo-supervised approaches overfit to label errors.

Method

The system utilizes the Large variant of WavLM (7-layer conv + 24 Transformer layers) as a pre-trained model. First, layer-wise Equal Error Rate (EER) is evaluated on validation trials to identify the top-K speaker-discriminative layers (layers 9, 10, and 11). Utterance representations from each selected layer are clustered independently using Infomap. A co-occurrence matrix and Hungarian matching align cluster IDs across layers, and only samples with identical cross-layer assignments are retained as high-confidence set H; remaining samples and small clusters form the unlabeled set L.

The student architecture consists of a PTM sub-network (using the first 10 layers of WavLM combined via learnable weighted sum) feeding an ECAPA-TDNN encoder and classifier. An Exponential Moving Average (EMA) teacher (decay α = 0.999) provides targets. The training loss combines supervised Label Noise Correction (LNC) loss on H (with time-dependent weight β_t increasing to 1.0) and unsupervised objectives on L. For L, two augmented views are generated; the teacher provides soft probability distributions with temperature T = 2.0, filtered by a confidence threshold τ = 0.6. Knowledge distillation loss (L_kd) is computed using KL divergence over confident indices, and a view-invariant embedding consistency loss (L_cons) is applied batch-wise, weighted by λ = 0.5.

After training, offline label correction splits utterances into chunks to filter ambiguous labels (discarding utterances where the majority label appears in under 50% of chunks), and the model is iteratively retrained for 4 total iterations, updating the pseudo-labels via clustering on embeddings from the current model.

Experimental setup

Trained on VoxCeleb2 dev (1,092,009 utterances from 5,994 speakers) using 2-second non-overlapping chunks with MUSAN and RIR data augmentation. Evaluated on VoxCeleb1-O using Equal Error Rate (EER) and minimum normalized detection cost function (MinDCF at P_target = 0.01). Implemented using AdamW with a cyclic learning rate schedule across a two-stage training recipe (10 epochs freezing PTM-sub, 5 epochs full fine-tuning with max learning rates 5e-4 and 1e-5). Compared against baselines including MBPC, AT-HT, CA-DINO, LGL, PTM-LC, and IPL.

Results

In Iteration-1, baseline S1 (PTM-LC equivalent using single-layer H1) achieves 2.64% EER. Adding HCPLS with K=3 layers (S3) improves EER to 2.40%. Incorporating DINO-style self-distillation on the low-confidence set (S5) further reduces EER to 2.10% (a 20.5% relative drop over S1) with a MinDCF of 0.212. Ablating the knowledge distillation loss (S7) causes a sharp degradation to 2.40% EER, demonstrating it is the primary driver over the consistency loss (S6, 2.17%). Across 4 iterative refinement cycles, the full method reaches 1.09% EER and 0.131 MinDCF on VoxCeleb1-O, outperforming prior arts like PTM-LC (1.25%) and IPL (1.14% in 11 iterations).

SystemTrain SetEER (%)MinDCF
PTM-LC [21]H12.55-
S1 (Baseline)H12.640.259
S3 (HCPLS K=3)H32.400.236
S5 (HCPLS + DINO-SD)H3 + L32.100.212
Proposed (Iter-4)Iterative1.090.131

Limitations

The evaluation is restricted to the VoxCeleb benchmark datasets, leaving multi-domain or cross-corpus robustness under-explored. The framework relies on a validation trial set to select top-K PTM layers, which introduces dependency on labeled validation data domain characteristics. Additionally, the multi-stage iterative pipeline with heavy pre-trained models demands substantial GPU compute resources.

Why read this

Read this if you work on self-supervised speaker recognition and want to learn how to effectively combine multi-layer pre-trained feature hierarchies with teacher-student self-distillation to handle noisy pseudo-labels.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Privacy-preserving speaker verification, forensic audio analysis, and large-scale speaker diarization pipelines using unlabeled speech.

Institutions

Malanshan Audio and Video Laboratory

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1965