TL;DR — LISE is a label-free post-hoc framework that decomposes pretrained speaker embeddings into a compact set of orthogonal, non-negative components, achieving human-verified perceptual interpretability (83.9% listening discrimination accuracy) with negligible automatic speaker verification performance degradation.
Key contributions
- Proposes LISE, an unsupervised, label-free post-hoc framework decomposing speaker embeddings into K orthogonal, non-negative continuous components.
- Preserves speaker verification performance on VoxCeleb1-O with minimal equal error rate (EER) degradation compared to original embeddings.
- Establishes a rigorous listening-test evaluation protocol to validate human perceptual discriminability of the decomposed components.
- Demonstrates robustness of the decomposition under data scarcity by training on reduced (75% and 50%) portions of the dataset.
Problem
Modern speaker embeddings from deep neural networks lack structured, perceptually verifiable explanations for their encoded vocal characteristics, while existing interpretability methods have severe flaws. Probing classifiers and attribute-disentanglement methods require expensive attribute labels, fail to capture fine-grained continuous vocal traits like breathiness, and cause severe ASV performance drops (e.g., EER jumping from 2.3% to 17.0%). Meanwhile, intrinsic sparse binary representations from text-inspired sparse autoencoders fail to align with how human listeners perceive continuous, low-dimensional human voice characteristics.
Method
LISE takes a pretrained speaker embedding (such as 512-dim x-vectors or 192-dim ECAPA-TDNN) and decomposes it post-hoc into components without retraining the backbone encoder. Given a component matrix and non-negative weights , the speaker representation is reconstructed via non-negative least squares coupled with an orthogonality regularization term using a hyperparameter .
The framework intentionally enforces three structural constraints: (1) low dimensionality (, optimized at ) to capture continuous voice variation rather than thousands of sparse binary activations; (2) non-negative additive weights () to eliminate ambiguous subtractive interactions and enable gradual perceptual scaling; and (3) orthogonality of the component matrix to reduce redundancy and ensure distinct, independent axes of variation.
The system is trained using 5,994 speaker embeddings from the VoxCeleb2 training set (approx. 1.1M utterances averaged per speaker) over 200 epochs on a single NVIDIA 2080Ti GPU. At inference, the learned component weights serve as an interpretable representation that can directly condition downstream models like SpeechT5 for controllable voice synthesis prototypes.
Experimental setup
Evaluated on the VoxCeleb1-O test set using Equal Error Rate (EER) via cosine similarity on reconstructed embeddings. Compared against three baselines: Principal Component Analysis (PCA), Luu et al. (attribute-supervised adversarial training), and Iben et al. (high-dimensional binary embeddings). Human perceptual evaluation involved 25 participants with self-reported normal hearing performing a 95-judgment discrimination task per method across 35 components.
Results
On VoxCeleb1-O, LISE achieves an EER of 3.08% for x-vectors (vs. 2.30% original, 3.02% PCA, 3.34% Iben et al., and 6.70% Luu et al.) and 2.10% for ECAPA-TDNN (vs. 1.80% original, 3.28% PCA, and 2.50% Iben et al.). Reducing the training data to 75% and 50% only slightly degrades EER to 3.13% and 3.23% for x-vectors, and 2.15% and 2.18% for ECAPA-TDNN, demonstrating high data efficiency.
In human listening evaluations, LISE achieves an overall accuracy of 83.9%, substantially outperforming PCA (59.1%) and Iben et al. (49.0%). Furthermore, 94.2% of LISE's components exceed a 70% listener accuracy threshold (compared to 11.4% for PCA and 0% for Iben et al.), and participant consistency remains high with individual listener accuracies tightly clustered between 73.7% and 91.6%.
| System / Condition | x-vector EER (%) ↓ | ECAPA-TDNN EER (%) ↓ | Listening Accuracy (%) ↑ |
|---|---|---|---|
| Original Embeddings | 2.30 | 1.80 | — |
| Luu et al. [12] | 6.70 | — | — |
| Iben et al. [14] | 3.34 | 2.50 | 49.0 |
| PCA | 3.02 | 3.28 | 59.1 |
| LISE (Ours, Full Data) | 3.08 | 2.10 | 83.9 |
| LISE (Ours, 50% Data) | 3.23 | 2.18 | — |
Limitations
The framework assumes English-dominant data (VoxCeleb), resulting in a small fraction of components (3 out of 35) capturing language patterns instead of speaker-intrinsic vocal features. The current evaluation relies on a post-hoc decomposition of static speaker-level averages rather than end-to-end training or real-time utterance-level temporal streaming. Subjective listener studies were constrained to 25 participants evaluating 35 components.
Why read this
Speech researchers and ML engineers looking to unpack opaque speaker embeddings into human-interpretable axes without sacrificing verification performance will find this a foundational read. It provides a concrete blueprint for combining non-negative continuous decomposition with rigorous perceptual listening validation.
Code
Applications
Controllable text-to-speech voice synthesis, model bias diagnosis, speaker embedding visualization, and auditory voice analysis.
Institutions
University of Southampton, Hong Kong Polytechnic University, University of Edinburgh
Funding / 經費: Engineering and Physical Sciences Research Council, National Edge AI Hub for Real Data: Edge Intelligence for Cyberdisturbances and Data Quality, Responsible AI UK
Related
- Do speech foundation models perceive speaker similarity as humans do? — same problem · relatedness 2.1/3
- Learning task-specific subspaces via interventional post-training of speech foundation models — same problem · relatedness 2.0/3
- Learning Multiple Utterance-Level Attribute Representations with a Unified Speech Encoder — same problem · relatedness 2.0/3
- Privacy and quality trade-off in real-time speaker anonymization via editing of age and sex attributes — complementary · relatedness 1.9/3
- Beyond task performance: Decoding bioacoustic embeddings with speech features — shared technique · relatedness 1.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-537