All papers
Speaker recognitionFull-paper digest

SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization

Petr Pálka, Jiangyu Han, Prachi Singh, Marc Delcroix, Naohiro Tawara, Lukáš Burget

Code & resourcesgithub.com/BUTSpeechFIT/DiariZen

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.1 KB · Ready to paste

Preview copied content

TL;DR — SphereVBx replaces the Gaussian PLDA backend in VBx with Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA), providing a principled Bayesian clustering framework for hyperspherical embeddings that simplifies end-to-end neural diarization with vector clustering (EEND-VC) while improving diarization error rate (DER). It achieves an average DER of 12.48% across eight benchmarks compared to the 12.65% baseline, while eliminating ad-hoc heuristic post-processing steps.

Key contributions

  • Introduces SphereVBx, substituting the Gaussian PLDA backend in VBx with T-PSDA to better model length-normalized hyperspherical embeddings trained with angular-margin objectives.
  • Proposes a parameter-free variant (SphereVBx-PF) based on cosine scoring that matches tuned backend performance without requiring pretrained T-PSDA model parameters.
  • Replaces heuristic short-embedding removal with continuous duration-based reliability weights that scale down the contribution of noisy segments during variational inference.
  • Develops a Multi-Stream (MS) variant (MS-SphereVBx) that natively enforces the intra-window cannot-link constraint directly inside the probabilistic inference loop via an outer-sum tensor formulation.

Problem

Modern speaker embeddings are length-normalized and optimized using angular-margin losses, making cosine similarity or spherical distributions theoretically more appropriate than traditional Gaussian PLDA models. However, standard VBx relies on Gaussian assumptions, forcing state-of-the-art EEND-VC systems to rely on heuristic steps such as filtering out unreliable short embeddings (<1.6s) and ad-hoc cosine-similarity reassignment. These decoupled pipelines complicate the architecture and hurt generalization, making a unified, geometry-aware probabilistic clustering framework necessary.

Method

SphereVBx reformulates the Variational Bayes clustering framework by substituting the Gaussian PLDA distributions with Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA), modeling embeddings on the unit hypersphere via von Mises–Fisher (vMF) mixtures. The speaker-specific directions live on the hypersphere, mapped through a subspace matrix K∈RD×dK \in \mathbb{R}^{D \times d} (with d=128d=128 dimensions projected from 256D inputs via LDA equivalent), where concentration parameters κw\kappa_w and κb\kappa_b control within- and between-speaker variances. In the parameter-free variant (SphereVBx-PF), setting d=Dd=D, κb=0\kappa_b=0, and κw=1\kappa_w=1 reduces the log-likelihood ratio scores to a monotonic function of cosine similarity without requiring training data or pretrained weights.

Instead of discarding short segments, duration-based reliability weights wtw_t are multiplied by posterior responsibilities (γts′=wtγts\gamma_{ts}' = w_t \gamma_{ts}) to smoothly dampen the impact of unreliable frames during variational updates. The inference alternates between estimating the vMF variational posterior q(ys)q(\mathbf{y}_s) and updating the posterior responsibilities γts\gamma_{ts}. To honor the EEND-VC intra-window 'cannot-link' constraint—where local speakers from the same 16-second window cannot share global identities—the Multi-Stream (MS-SphereVBx) variant builds an LL-way tensor via an outer sum of row-wise log-scores. Invalid joint assignments containing repeated global speaker indices are pruned, and marginalizing the valid normalized tensor yields the speaker-level responsibilities required for the standard vMF variational update equations.

Experimental setup

Evaluated across eight standard diarization benchmarks: AMI (far-field ch1), AISHELL-4, AliMeeting (far-field ch1), NOTSOFAR-1, MSDWild, DIHARD3 full, RAMC, and VoxConverse. Systems are compared against standard VBx cascades and the state-of-the-art DiariZen EEND-VC baseline built on a WavLM Large feature extractor and Conformer encoder. Performance metrics include Diarization Error Rate (DER) without a forgiving collar, macro-averaged DER, and Mean Speaker Count Error (MSCE).

Results

In the cascaded VAD+VBx+OSD pipeline, SphereVBx achieves an average DER of 22.1% across datasets (improving upon the 22.7% re-evaluated baseline and 22.6% original VBx), while its parameter-free variant SphereVBx-PF scores 22.2% DER. Within the EEND-VC framework, the baseline DiariZen system records an average DER of 12.65% with an MSCE of 0.37, whereas standard SphereVBx improves average DER to 12.52% (MSCE 0.34). The fully-constrained MS-SphereVBx-PF model achieves the best average DER of 12.48% (MSCE 0.37), matching or beating concurrent published SOTA while requiring no pretrained backend parameters.

SystemAISAliMAMIDH3MSDNSFRAMCVoxCAverage DER (%)
Baseline [25]9.910.813.914.515.716.711.08.812.65
SphereVBx9.610.713.714.315.516.710.98.812.52
MS-SphereVBx9.610.713.714.215.616.610.98.912.52
MS-SphereVBx-PF9.710.613.714.115.816.510.68.912.48

Limitations

The current reliability weighting strategy relies entirely on a simple binary threshold derived from segment duration (1.6s) rather than continuous confidence metrics extracted from acoustic features. Multi-stream combinatorial tensor expansions can become computationally expensive for windows with very high local speaker counts (LL). The evaluation is restricted to speaker diarization benchmarks, leaving the integration with face recognition or spatial audio source separation for future work.

Why read this

Speech engineers and ML researchers working on speaker diarization or vector clustering should read this paper to learn how to replace heuristic post-processing rules and Gaussian PLDA backends with a unified, geometry-aware spherical Bayesian framework.

Code

Applications

Speaker diarization pipelines, multi-speaker conversational transcription systems, and acoustic meeting analysis tools.

Institutions

Brno University of Technology, NTT

Funding / 經費: European Union, Czech Ministry of Education, Youth and Sports

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2224