All papers
Applications & otherFull-paper digest

Causal Redundancy in Speech Representations: The Hydra Effect and Limits of Sparse Disentanglement in WavLM

Patalee Narasinghe, Kasindu Bandara, Vilash Nawagamuwa, Janak Senevirathne, Uthayasanker Thayasivam

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.5 KB · Ready to paste

Preview copied content

TL;DR — This paper investigates internal feature encoding in WavLM and discovers a "Hydra effect" where discrete neuron or sparse autoencoder latent ablations fail to suppress acoustic features due to massive representational redundancy. To achieve targeted feature removal, the authors show that Iterative Null-space Projection (INLP) must be used to erase entire continuous linear subspaces.

Key contributions

  • Demonstrates that raw WavLM early-layer neurons exhibit severe polysemanticity across low-level acoustic descriptors (GeMAPS).
  • Applies JumpReLU Sparse Autoencoders (SAEs) to WavLM representations and identifies the "Hydra effect"—where zeroing out top-ranked interpretable latents fails to suppress acoustic features.
  • Proves that continuous acoustic features in speech SSL models reside in distributed linear subspaces rather than isolated monosemantic neurons.
  • Applies Iterative Null-space Projection (INLP) to successfully achieve targeted subspace erasure, showing that loudness occupies an isolated subspace while spectral features exhibit heavy mutual coupling.

Problem

Self-supervised speech models like WavLM build hierarchical representations transitioning from low-level acoustics to high-level semantics, yet their internal mechanisms remain opaque. Prior interpretability methods rely heavily on linear probing, which is merely correlational, or attempt neuron-level feature disentanglement via Sparse Autoencoders (successful in NLP). However, the continuous, highly redundant nature of speech signals creates an interpretability gap: it is unknown whether acoustic features are encoded in isolated neurons or distributed manifolds, and whether standard neuron-level interventions can causally manipulate them.

Method

The authors analyze WavLM-Base+ (94M parameters, 12 transformer layers, 768-dimensional hidden states, 12 attention heads), pretrained on 60k hours of English speech. They extract frame-wise hidden states from Layer 1 and pair them via linear interpolation with six openSMILE-derived GeMAPS descriptors (pitch, F1 frequency, Alpha Ratio, Hammarberg Index, HNR, Loudness) sampled at 10 ms vs WavLM's 20 ms stride.

To address polysemanticity, they train a JumpReLU Sparse Autoencoder mapping residual activations from R768\mathbb{R}^{768} to a high-dimensional sparse latent space R8192\mathbb{R}^{8192} (expansion factor α≈10.67\alpha \approx 10.67). The JumpReLU SAE optimizes reconstruction error alongside an L0L_0 sparsity penalty controlled by learned per-feature thresholds θi\theta_i and the Heaviside step function, avoiding L1L_1 magnitude shrinkage. The SAE is trained on 1M voice frames.

For causal analysis, the authors perform two types of intervention. First, they rank SAE latents using Lasso probe SHAP values and systematically ablate the top-N most critical latents to test for the Hydra effect. Second, to bypass this redundancy, they employ Iterative Null-space Projection (INLP) to learn linear predictors for specific features and project representations onto their null spaces, completely erasing the corresponding linear subspaces while checking for cross-feature selectivity via R2R^2 retention.

Experimental setup

Evaluated on a multi-corpus, speaker-independent dataset combining RAVDESS, CREMA-D, and TIMIT (~6,000 utterances, ~10 hours of audio, 3s average duration). Baselines include random label shuffling (for probing) and random projection controls (for INLP). Metrics include coefficient of determination (R2R^2) for continuous probes, SAE reconstruction MSE, explained variance (R2R^2), and L0L_0 sparsity. All experiments are run on a single NVIDIA RTX 4070 GPU with 16 GB memory.

Results

WavLM Layer 1 linear probes achieve high R2R^2 for acoustic features like F0F_0 and loudness, confirming linear encoding in early layers. The JumpReLU SAE achieves an R2R^2 reconstruction fidelity of 0.860 (MSE = 0.0069 vs mean baseline 0.0493) with a strict sparse bottleneck of mean L0=97.8L_0 = 97.8 active latents (~1.2% utilization). Despite this sparsity, iterative ablation of the top 50 to 100 SAE latents shows that pitch (F0F_0) and alpha ratio retain >98% of their probe R2R^2 (Ret. @ 50: 98.85% for pitch SAE, 98.31% for WavLM L1), proving the Hydra effect where discrete ablation fails.

In contrast, INLP subspace erasure successfully suppresses target features by >94% (F0F_0 drops 94.41%, loudness drops 94.66%, HNR drops 95.35%). Off-diagonal analysis reveals that loudness erasure is highly selective (<1.5% cross-drop on other probes), whereas spectral features (alphaRatio, hammarbergIndex, mfcc1) exhibit strong mutual coupling with cross-drops exceeding 75%, indicating a shared subspace.

Feature & SpaceRet. @ 50NcriticalN_{critical}
Pitch (F0F_0) - WavLM L198.31%> 500
Pitch (F0F_0) - SAE Latent98.85%1000
AlphaRatio - WavLM L198.42%> 500
AlphaRatio - SAE Latent98.98%1000
Loudness - WavLM L196.96%500
Loudness - SAE Latent97.07%500

Limitations

The study is restricted to WavLM-Base+ and a relatively small 10-hour English subset combining RAVDESS, CREMA-D, and TIMIT, limiting generalization to larger speech models (e.g., Large/XL variants) and multilingual or noisy settings. The evaluation focuses exclusively on low-level continuous GeMAPS acoustic descriptors extracted from early layers, leaving higher-level semantic, linguistic, and speaker characteristics unexamined via subspace erasure. Additionally, INLP is limited to linear subspace removal, meaning potential non-linear distributed redundancy remains unmeasured.

Why read this

Speech ML researchers and interpretability engineers studying self-supervised models should read this to understand why applying text-domain Sparse Autoencoders out-of-the-box fails to cleanly isolate speech features due to distributed redundancy (the Hydra effect). It provides a methodological roadmap demonstrating that subspace-level interventions like INLP are necessary for causal manipulation of continuous acoustic representations.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Targeted editing, debiasing, or disentangling of continuous acoustic attributes in speech generation, voice conversion, and speech representation learning pipelines.

Institutions

University of Moratuwa

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3316