All papers
Speaker recognitionFull-paper digest

Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge

Ze Li, Xiaoxiao Miao, Juan Liu, Ming Li

Code & resourcesgithub.com/ZXHY-82/LI-MSV-TidyVoice2026

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.3 KB · Ready to paste

Preview copied content

TL;DR — A language-invariant multilingual speaker verification system is developed for the TidyVoice 2026 Challenge by combining a w2v-BERT 2.0 backbone, SphereFace2 loss, language-adversarial training via gradient reversal, and zero-shot TTS augmentation, achieving an EER of 0.89% on the development set.

Key contributions

  • Adapts the 4.5-million-hour multilingual self-supervised w2v-BERT 2.0 model with Layer Adapters, Multi-scale Feature Aggregation (MFA), and Attentive Statistics Pooling for robust speaker embedding extraction.
  • Applies a Gradient Reversal Layer (GRL) and an auxiliary language classifier for language-adversarial training to force embeddings to be language-invariant.
  • Investigates the use of SphereFace2 (configs A and C) over ArcFace, aligning the training loss's hyperspherical binary classification objective with pairwise verification protocols.
  • Explores multilingual zero-shot text-to-speech (ZS-TTS) data augmentation using Qwen3-TTS-12Hz-1.7B-Base across 10 languages to synthesize 349,500 utterances from 3,495 reference files.

Problem

Multilingual speaker verification suffers significant performance degradation under language mismatch conditions, worsened by an over-reliance on English-centric corpora. Limited multilingual speech per speaker causes embeddings to entangle speaker identity with language-specific traits, reducing cross-lingual generalization. Prior standard multi-class formulations like ArcFace also mismatch the pairwise similarity evaluation protocol used in verification.

Method

The system processes 80-dimensional fbank features (25ms window, 10ms hop) through a 24-layer Conformer w2v-BERT 2.0 backbone. Hidden representations from every layer pass through individual Layer Adapters to lower dimensionality and adapt domains. Adapted features are concatenated, aggregated via Attentive Statistics Pooling (ASP), and projected to yield a speaker embedding. Low-Rank Adaptation (LoRA) is used during fine-tuning.

To decouple speaker identity from language, an auxiliary language classifier is attached to the embedding space via a Gradient Reversal Layer (GRL). The total objective combines speaker loss (ArcFace or SphereFace2) and scaled language loss: L = L_spk(C_spk(e), y_spk) + lambda_lang * L_lang(C_lang(GRL_lambda_GRL(e)), y_lang), where lambda_GRL and lambda_lang are both set to 0.1. Training occurs in two stages: pre-training on large public datasets (VoxCeleb2, VoxBlink2, 3D-Speaker, KeSpeech, CN-Celeb1&2) using an AdamW optimizer (learning rate 1e-4 decaying to 1e-5 via StepLR, batch length 200-300 frames) with on-the-fly MUSAN noise and RIR reverberation, followed by domain adaptation on TidyVoiceX.

Score calibration utilizes a Quality Measure Function (QMF) driven by logistic regression over trial metadata (enrollment/test durations, embedding magnitudes, and max-min normalized SNRs). For synthetic augmentation, Qwen3-TTS-12Hz-1.7B-Base translates LibriTTS English text via M2M100 into 10 target languages, using reference audio transcribed by Whisper-large-v3.

Experimental setup

Evaluated on the TidyVoice 2026 Challenge development set (tv26-dev) and evaluation subsets (tv26-eval-A with seen languages, tv26-eval-U with 38 unseen languages). Training uses VoxCeleb2, VoxBlink2, 3D-Speaker, KeSpeech, CN-Celeb1&2, and TidyVoiceX. Baselines include the official TidyVoice 2026 challenge baseline (SimAMResNet34). Metrics reported are Equal Error Rate (EER %) and minimum Detection Cost Function (mDCF at p_target = 0.01).

Results

Fine-tuning w2v-BERT 2.0 without TidyVoiceX data yields an EER of 2.74% on tv26-dev, an 11% relative reduction over the official baseline's 3.07%. Replacing ArcFace with SphereFace2-C drastically improves performance, dropping the tv26-dev EER to 1.065% (and 0.893% when incorporating CN-Celeb1&2 and QMF score calibration). On the challenge evaluation splits, the full system achieves 2.458% EER on tv26-eval-A and 4.451% EER on tv26-eval-U.

Ablations demonstrate that fine-tuning solely on TidyVoiceX specializes well for seen languages (tv26-dev/eval-A) but hurts unseen language generalization (tv26-eval-U), which requires pre-training data mixing. Synthetic speech augmentation via Qwen3-TTS achieves a competitive 1.022% EER on tv26-dev when trained entirely on synthetic data (using only 1/10th the real reference data), but combining it with real data yields no further gains under sufficient data regimes, indicating a domain mismatch.

Systemtv26-dev EER (%)tv26-dev mDCFtv26-eval-A EER (%)tv26-eval-U EER (%)
Official Baseline [22]3.070.829.05811.59
w2v-BERT 2.0 Based (ArcFace)1.4660.66--
w2v-BERT 2.0 Based SF2-C1.0650.623.0614.338
w2v-BERT 2.0 + GRL + KeSpeech0.9370.602.9645.020
w2v-BERT 2.0 + GRL + QMF + CN-Celeb0.8930.602.4584.451

Limitations

Synthetic speech augmentation fails to improve performance when large amounts of real training data are already accessible due to domain gaps. The language-adversarial training provides only modest improvements in suppressing language-specific variations while occasionally hurting zero-shot unseen language error rates. Evaluation is constrained to the specific language distributions provided by the TidyVoice 2026 Challenge splits.

Why read this

Researchers building state-of-the-art multilingual speaker verification systems will learn how to effectively combine large self-supervised conformer models, hyperspherical binary classification losses, and gradient-reversal language invariance strategies.

Code

Applications

Cross-lingual speaker verification, multilingual speaker recognition, and zero-shot voice cloning security systems.

Institutions

Wuhan University, Chinese University of Hong Kong, Shenzhen, Duke Kunshan University

Funding / 經費: National Natural Science Foundation of China, Yangtze River Delta Science and Technology Innovation Community Joint Research Project

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2437