TL;DR — The paper introduces Continuous Influence Permutation Invariant Training (CI-PIT) to resolve hierarchical permutation inconsistency in self-conditioned end-to-end speaker diarization, achieving a lower Diarization Error Rate (DER) of 7.52% on CALLHOME 2-speaker compared to 8.30% for the baseline.
Key contributions
- Formalized Hierarchical Permutation Inconsistency (HPI), showing how independent per-layer PIT in self-conditioned EEND injects structured permutation noise into the conditioning path.
- Proposed Group-Wise PIT (GW-PIT) to enforce a hard global permutation constraint across all intermediate layers.
- Proposed Continuous Influence PIT (CI-PIT), applying Gaussian-weighted cost matrix aggregation to softly regularize inter-layer permutations while preserving local representation flexibility.
- Introduced three diagnostic metrics—Permutation Mismatch Ratio (PMR), Adjacent Switching Ratio (ASR), and Epoch Switching Ratio (ESR)—to quantify permutation dynamics across layers and epochs.
Problem
Self-conditioned end-to-end neural diarization (EEND) improves performance by iteratively feeding intermediate speaker predictions back into subsequent encoder layers for progressive refinement. However, conventional layer-wise permutation invariant training (LW-PIT) independently resolves the optimal speaker permutation at each layer via hard minimization. In a self-conditioned architecture, these uncoordinated per-layer permutations cause adjacent layers to adopt conflicting speaker assignments, a failure mode termed Hierarchical Permutation Inconsistency (HPI). HPI propagates permutation noise directly into the self-conditioning path, destroying the progressive feature refinement intended by the architecture and destabilizing training convergence.
Method
The base architecture uses an 8-layer Transformer encoder (256-dim embedding, 4 attention heads, 2048-dim FFN) with non-autoregressive (NA) multi-head self-attention extracting speaker attractors via 500-frame input segments. Intermediate predictions are fed back into subsequent layers using a shared weight matrix . To fix HPI, the paper replaces standard Layer-Wise PIT () with two alternatives.
Group-Wise PIT () relocates the minimum operator outside the summation over layers, enforcing a single global permutation across all encoder layers simultaneously. While this completely eliminates inter-layer mismatch by construction, it imposes an overly rigid constraint that restricts layer-specific gradients.
Continuous Influence PIT () provides a soft regularization scheme through a three-step process: (1) computing Gaussian weights to determine the influence of layer on target layer ; (2) aggregating pairwise BCE cost matrices from neighboring layers into a unified cost matrix ; and (3) solving for the optimal permutation via the Hungarian algorithm on . By tuning (set to 1.0 optimally), CI-PIT interpolates between local flexibility and global consistency, suppressing permutation noise while maintaining necessary layer-wise differentiation.
Experimental setup
Models were trained on simulated conversations (SimConv) containing 2,478 hours for 2-speaker and 2,513 hours for 3-speaker scenarios, generated using LibriSpeech source speech augmented with MUSAN noise and downsampled to 8 kHz. Evaluation was conducted on the CALLHOME (NIST SRE 2000) dataset (adaptation/test splits with lengths of ~3 hours per condition). The system uses the EEND-NA architecture trained from scratch for 100 epochs on SimConv with the Noam scheduler, then fine-tuned on CALLHOME adaptation for 100 epochs at a learning rate of using Adam optimization. Diarization Error Rate (DER) with a 0.25-second forgiveness collar and an 11-frame median filter post-processing is the primary metric.
Results
On the CALLHOME 2-speaker test set, the proposed () achieves a headline DER of 7.52%, outperforming the baseline (8.30%) and (8.02%), primarily driven by a drop in False Alarm rate from 3.65% to 2.59%. On the more challenging 3-speaker scenario, achieves the best DER of 12.52% compared to the baseline's 13.04% and 's 13.66%, demonstrating that the hard global constraint of GW-PIT degrades when scaling to more speakers whereas CI-PIT scales effectively. Ablations isolating self-conditioning (SC) show that SC alone yields negligible gains with LW-PIT (8.34% to 8.30%) due to noise corruption, whereas pairing SC with CI-PIT yields a substantial drop from 8.67% to 7.52% DER. Varying the Gaussian width demonstrates that broadening the scope degrades performance (7.97% at and 8.12% at ), confirming that localized, layer-proximate alignment is critical.
| Training Objective | MI (%) | FA (%) | CF (%) | DER (2-spk %) | DER (3-spk %) |
|---|---|---|---|---|---|
| (Baseline) | 4.11 | 3.65 | 0.53 | 8.30 | 13.04 |
| 4.21 | 3.17 | 0.64 | 8.02 | 13.66 | |
| () | 4.36 | 2.59 | 0.57 | 7.52 | 12.52 |
Limitations
The evaluation is restricted to telephone speech (CALLHOME and LibriSpeech-based simulations downsampled to 8 kHz), leaving open how the method performs on wideband 16kHz audio or reverberant environments requiring Room Impulse Response (RIR) augmentation. The study focuses exclusively on 2- and 3-speaker scenarios, so scaling behavior for highly overlapping or many-speaker conversations (e.g., meeting room domains) remains unexplored. Furthermore, compute overhead introduced by the Hungarian algorithm cost matrix aggregation across layers during training is not explicitly quantified.
Why read this
Researchers and engineers building self-conditioned end-to-end neural diarization systems should read this paper to understand how layer-wise permutation inconsistency destabilizes training, and adopt CI-PIT as a plug-and-play training objective to properly unlock the benefits of self-conditioning.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Speaker diarization systems for multi-speaker conversational transcription, meeting transcription, and call center analytics.
Institutions
Incheon National University
Funding / 經費: KOITA, MSIT
Related
- Two-Level Uncertainty Suppression for Robust Meeting Diarization — same problem · relatedness 2.7/3
- Speaker Separation via Audio Language Modeling — same problem · relatedness 2.6/3
- Bidirectional Retention Network-based Segmentation Model for Speaker Diarization — same problem · relatedness 2.6/3
- SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization — same problem · relatedness 2.6/3
- Multi-Speaker Embeddings With Weakly Supervised Speaker Activity Detection For Granular Speaker Diarization — same problem · relatedness 2.4/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1198