All papers
Speech synthesisFull-paper digest

Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity

Ismail Rasim Ulgen, John Hansen, Carlos Busso, Berrak Sisman

Code & resourceschoughtotem.github.io/subcentervc_demo

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.9 KB · Ready to paste

Preview copied content

TL;DR — This paper proposes a sub-center modeling framework for speaker embeddings to bridge the gap between recognition-driven compactness and generation-oriented variability, improving naturalness and pitch diversity in zero-shot voice conversion.

Key contributions

  • A generation-oriented perspective that reframes intra-speaker variance as a design parameter rather than noise to suppress.
  • A sub-center training formulation using an extended AAM-Softmax objective that learns multiple prototypes per speaker while maintaining discriminability.
  • Empirical demonstration via voice conversion that increased intra-speaker embedding variance improves naturalness, intelligibility, and prosodic expressiveness.
  • Analysis of temperature-scaling effects on sub-center utilization and the trade-off between speaker similarity and intelligibility.

Problem

Conventional speaker embeddings are trained for speaker recognition using large-scale classification objectives that enforce small intra-speaker variance and high inter-speaker separation. While this single-center prototype structure optimizes recognition performance, it collapses acoustic variations caused by prosody, style, and emotion. Consequently, these recognition-optimized embeddings are suboptimal when used as conditioning signals for downstream speech generation systems like text-to-speech and voice conversion, which require structured intra-speaker diversity to sound natural and expressive.

Method

The authors modify the standard ECAPA-TDNN architecture by replacing the single class-center classification head with a sub-center formulation. The classifier weight matrix is expanded to Wc∈RL×N×CW_c \in \mathbb{R}^{L \times N \times C}, maintaining CC sub-centers per speaker class (nn), where LL is the embedding dimension and NN is the number of speakers. The similarity sn,cs_{n,c} between an embedding vector xix_i and the cc-th sub-center of class nn is aggregated using temperature-scaled softmax weights to compute a sub-class similarity score s~n\tilde{s}_n, which is then fed into an adapted Additive Angular Margin Softmax (AAM-Softmax) loss with margin mm and scale kk.

To evaluate the embeddings in a generative context, the authors adopt a speech-resynthesis voice conversion pipeline built on discrete HuBERT units (from the 6th layer clustered via k-means with K=100K=100 on LibriSpeech-clean-100) and pitch units extracted using the Dio algorithm combined with a VQ-VAE. A HiFi-GAN decoder generates expressive waveforms conditioned on these linguistic units, pitch units, and a 192-dimensional sub-center speaker embedding extracted from a reference utterance. The embedding network is trained on VoxCeleb2 using the Adam optimizer with a base learning rate of 1e−41e-4, a cyclic learning rate schedule, online data augmentation (noise and reverberation), batch size 32, margin m=0.2m=0.2, and scale s=30s=30.

Experimental setup

Speaker embeddings are trained on VoxCeleb2. Voice conversion experiments use the VCTK corpus (110 English speakers, ~400 utterances each), with 90 speakers for training and 20 unseen speakers for zero-shot evaluation. Baselines include standard ECAPA-TDNN embeddings. Evaluations use Equal Error Rate (EER) on VCTK and VoxCeleb1-E test sets, Word Error Rate (WER) and Character Error Rate (CER) via a Wav2Vec2 ASR model, Speaker Embedding Cosine Similarity (SECS), and subjective listening tests including MOS for naturalness, SMOS for speaker similarity, and ABX tests for prosody.

Results

Sub-center ECAPA-TDNN models achieve lower or comparable EER on speaker verification while successfully increasing intra-to-inter-class variance ratios. For instance, the C=20C=20 configuration yields an intra/inter-class variance ratio of 0.47 on VCTK (vs 0.42 for baseline ECAPA-TDNN) and 0.91 on VoxCeleb1-E (vs 0.66 for baseline), alongside an improved VCTK verification EER of 1.55%. In voice conversion objective evaluations, the C=20C=20 sub-center model achieves the lowest WER (13.93% vs 14.84% baseline) and CER (6.41% vs 6.82% baseline), whereas lower-temperature variants (T=0.1T=0.1) optimize SECS (65.86% vs 64.04%). Subjectively, the C=20C=20 model significantly outperforms the baseline in naturalness MOS (3.18 vs 2.94) and similarity MOS (2.88 vs 2.65), and wins clear preference in ABX prosody tests.

SystemEER (%) ↓WER ↓SECS (%) ↑Naturalness MOS ↑
Baseline ECAPA-TDNN1.7114.8464.042.94
Sub-center (C=10,T=0.1C=10, T=0.1)1.4714.3265.862.89
Sub-center (C=10C=10)1.5014.6564.14—
Sub-center (C=20C=20)1.5513.9364.593.18

Limitations

The evaluation is restricted to English using a single voice conversion framework (VCTK and VoxCeleb datasets) and does not test cross-lingual or highly expressive pathological speech variations. The exploration of sub-center scaling is limited to C=10C=10 and C=20C=20 with two temperature values, leaving optimal hyperparameter tuning across diverse model capacities unexhausted. Additionally, the approach demonstrates an empirical trade-off where higher variance improves intelligibility and prosody at a minor cost to strict speaker identity similarity.

Why read this

Speech and ML researchers building personalization or zero-shot generation systems should read this to understand why standard recognition-trained speaker embeddings limit expressive output and how simple sub-center classification modifications resolve this bottleneck.

Code

Applications

Zero-shot voice conversion, expressive text-to-speech, and personalized multi-speaker audio generation systems.

Institutions

Johns Hopkins University, University of Texas at Dallas, Carnegie Mellon University

Funding / 經費: National Science Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-942