All papers
Deepfakes & securityFull-paper digest

Imperceptible Voiceprint Protection via Human-Machine Perception Discrepancy Feature Disentanglement

Chenlong Xue, Meng Sun, Qiang Zhang, Xiongwei Zhang, Kunyuan Li, Yuan Liao, Xiaoyi Ge, Kui Yao

Code & resourcescero529.github.io/voiceprint-protection-demo

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.3 KB · Ready to paste

Preview copied content

TL;DR — A two-stage adversarial protection framework isolates speaker identity from linguistic content via an information bottleneck and injects imperceptible perturbations into the speaker embedding space, achieving an 87.2% defense success rate against voice cloning while maintaining high audio quality (MOS 4.18).

Key contributions

  • A two-stage disentanglement-reconstruction framework that separates content codes and speaker embeddings to prevent perturbation leakage into linguistic features.
  • A speaker-embedding-space perturbation generator combined with time-frequency psychoacoustic masking constraints to target machine sensitivity without human perceptual degradation.
  • Superior white-box defense success rate (87.2% at tau=0.5) and robust black-box transferability across unseen voice cloning architectures (YourTTS, VALL-E, AdaptVC).

Problem

Existing voice cloning defenses suffer from a strict trade-off: waveform-level and frequency-domain signal perturbations either introduce audible artifacts or get stripped by audio preprocessing, while unconstrained embedding-space perturbations entangle with content features and ruin speech intelligibility. Prior methods also fail to generalize well when transferred to unseen black-box cloning systems. Solving this is critical to protect individuals from unauthorized voice cloning, financial fraud, and identity spoofing.

Method

The framework operates in two distinct stages. Stage 1 trains an AutoVC-based disentanglement network using a content encoder (3 convolutional layers plus a bidirectional LSTM with a bottleneck dimension of 32) and a speaker encoder (finetuned GE2E) to decompose mel-spectrograms (x∈RT×80x \in \mathbb{R}^{T \times 80}) into downsampled content codes (c∈RT′×32c \in \mathbb{R}^{T' \times 32}) and speaker embeddings (h∈R256h \in \mathbb{R}^{256}). An adversarial entropy maximization loss (LdisL_{dis}) forces content codes to yield a uniform speaker classification distribution, preventing speaker identity leakage.

Stage 2 freezes the Stage 1 network and trains a generator G(⋅)G(\cdot) to produce perturbations δ=G(h)\delta = G(h), yielding a protected embedding hadv=h+α⋅δh_{adv} = h + \alpha \cdot \delta (with scaling factor α=0.10\alpha = 0.10). The protected mel-spectrogram xadv=D(c,hadv)x_{adv} = D(c, h_{adv}) is synthesized into audio via a HiFi-GAN vocoder. The multi-objective generator loss comprises a mel-spectrogram loss (LmelL_{mel}), a psychoacoustic masking loss (LmaskL_{mask}) constrained against the MPEG psychoacoustic threshold θx(k)\theta_x(k), an LSGAN realism loss (LGANL_{GAN}), and a defense loss (LdeL_{de}) that minimizes cosine similarity between speech cloned from the stolen embedding and the original speaker.

Key hyperparameters include training Stage 1 for 500k iterations (λcd=1.0,λdis=0.01\lambda_{cd} = 1.0, \lambda_{dis} = 0.01) and Stage 2 for 200k iterations (α=0.10,λper=15.0,λde=5.0,λmask=0.1\alpha = 0.10, \lambda_{per} = 15.0, \lambda_{de} = 5.0, \lambda_{mask} = 0.1). This design ensures perturbations strictly target machine-vulnerable subspaces while remaining imperceptible to human ears.

Experimental setup

Evaluated on the VCTK corpus containing 110 English speakers resampled to 16kHz (80 for training, 10 for validation, 20 for testing, with 100 utterances per test speaker). Compared against time-domain baselines (Gaussian Noise, Voice Guard, CloneShield), frequency-domain baselines (VocalCrypt), and embedding-space baselines (MI-FGSM, RoVo). Metrics include Simprot (identity preservation), Simclone (cloning similarity), Defense Success Rate (DSR) at thresholds τ∈{0.3,0.4,0.5}\tau \in \{0.3, 0.4, 0.5\}, Mean Opinion Score (MOS) rated by 50 listeners, and Word Error Rate (WER) using Whisper-large.

Results

The proposed method achieves a white-box DSR of 87.2% at τ=0.5\tau = 0.5, outperforming RoVo (79.2%) and Voice Guard (79.8%), while maintaining a high MOS of 4.18 and a low WER of 5.30%. Simprot reaches 0.95 and Simclone drops to 0.13. Ablations show that removing the defense loss (LdeL_{de}) causes DSR to collapse to 12.5%, whereas removing the psychoacoustic masking loss (LmaskL_{mask}) drops MOS from 4.18 to 3.38 and raises WER to 8.92%, confirming the necessity of psychoacoustic constraints. In black-box cross-architecture evaluations, the method maintains superior DSR across YourTTS, VALL-E, and AdaptVC compared to signal-level and baseline embedding defenses.

MethodDomainSimprot ↑\uparrowSimclone ↓\downarrowDSR(%) τ=0.5\tau=0.5 ↑\uparrowMOS ↑\uparrowWER(%) ↓\downarrow
Original–1.000.910.04.534.78
Voice GuardTime0.870.1879.83.5210.25
VocalCryptFreq.0.880.2174.63.588.95
RoVoEmb.0.890.1879.23.727.15
OursEmb.0.950.1387.24.185.30

Limitations

The evaluation is restricted to clean English speech from the VCTK corpus and evaluated primarily against specific open-source cloning architectures (AutoVC, YourTTS, VALL-E, AdaptVC). The paper does not evaluate robustness against adaptive attacks where the attacker knows the exact defense mechanism, nor does it test performance under real-world acoustic degradations like lossy audio compression or environmental noise.

Why read this

Speech and security researchers looking to balance adversarial defense efficacy and perceptual transparency should read this to see how feature disentanglement and psychoacoustic constraints solve embedding-space vulnerability trade-offs.

Code

Applications

Voice data privacy protection, pre-emptive defense against zero-shot voice cloning, and anti-spoofing watermarking for personal audio.

Institutions

Army Engineering University of PLA, Chinese University of Hong Kong, Information Support Force Engineering University

Funding / 經費: National Natural Science Foundation of China, Natural Science Foundation of Jiangsu Province, China Postdoctoral Science Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-105