All papers
Deepfakes & securityFull-paper digest

Task-Aware Joint Pruning and Distillation for Efficient Audio Deepfake Detection

Miao He, Peng Cheng, Zhongjie Ba, Qing Wen, Li Lu, Xin Yang, Kui Ren

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

14.1 KB · Ready to paste

Preview copied content

TL;DR — This paper introduces a task-aware joint pruning and distillation framework to compress large self-supervised speech models for on-device deepfake detection, achieving a 10x parameter reduction (31.9M) with only a 1.30% average performance drop.

Key contributions

  • Proposes a cross-domain knowledge distillation strategy utilizing unlabeled out-of-domain data with various acoustic and codec variations to prevent compression-induced generalization loss.
  • Introduces movement-guided structured pruning that leverages fine-tuning weight movement scores as a structural importance prior to safeguard forgery-critical sub-networks.
  • Applies CKA (Centered Kernel Alignment) layer similarity analysis to identify optimal multi-level representative layers for hierarchical knowledge transfer across Transformer blocks.
  • Demonstrates aggressive compression of the 317M parameter XLSR-AASIST model down to 31.9M parameters and 23.3G FLOPs (90% sparsity) while preserving out-of-domain robustness.

Problem

State-of-the-art audio deepfake detectors rely heavily on self-supervised learning (SSL) backbones exceeding 300 million parameters, restricting them to cloud-only inference which introduces network latency and serious user privacy risks. Existing compression techniques are predominantly engineered for content-centric tasks like automatic speech recognition or speaker verification, leading to catastrophic out-of-domain performance drops and poor generalization when directly applied to forgery detection. Furthermore, existing compressed models degrade severely at high sparsity levels (>75%) due to the loss of subtle artifact-capturing capabilities.

Method

The framework operates in three main stages: task-specific preparation, joint pruning and distillation, and backend optimization. First, the pretrained XLS-R (0.3B) model is fine-tuned on ASVspoof2019 LA, during which structural movement scores are calculated as the L1 norm of the weight changes between initialization and fine-tuning (θft−θinit\theta^{ft} - \theta^{init}). A penalty regularization term with binary retention masks is then minimized to prune low-movement units while retaining critical structures in Multi-Head Attention (MHA) layers, FFN intermediates, and CNN channels.

Concurrently, multi-level cross-domain knowledge distillation is applied using unlabeled out-of-domain datasets (ASVspoof2021 LA, 21 DF, and In-the-Wild) to expose the student model to diverse codec and environmental distortions. Guided by Centered Kernel Alignment (CKA) layer-wise similarity analysis—which reveals three block-like processing stages—representative layers (specifically layers 5, 14, and 24) are chosen for feature matching using MSE or cosine distance losses. The final constrained multi-objective optimization objective enforces target sparsity via the Augmented Lagrangian method with Lagrangian multipliers, after which the pruned student model is jointly fine-tuned with the AASIST backend classifier.

Experimental setup

The method compresses the XLS-R (0.3B) frontend within the XLSR-AASIST baseline architecture. Datasets used include ASVspoof2019 LA for fine-tuning, ASVspoof2019 LA train, 2021 LA, 2021 DF, and In-the-Wild (unlabeled) for distillation, and evaluations span ASVspoof2019 LA (Dev/Eval), ASVspoof2021 LA, ASVspoof2021 DF, In-the-Wild, ASVspoof5, and FoR. Baselines include HJ-Pruning, Finetune-Pruning, and Hybrid-Pruning adapted to deepfake detection. The student is trained for 50k steps with a 5k warmup for target sparsity, learning rates of 2e-4 for model parameters and 2e-2 for auxiliary parameters, followed by 25k steps of secondary distillation.

Results

At 75% sparsity, the proposed framework achieves an EER of 8.43% on In-the-Wild and 17.75% on ASVspoof5, outperforming the second-best baseline (Finetune-Pruning) by 67% and 32% respectively. Remarkably, at 75% sparsity, the compressed model outperforms the uncompressed dense baseline on unseen datasets like ASVspoof5 (17.75% vs 19.92% EER) and FoR (10.51% vs 14.89% EER), demonstrating that structural pruning acts as a regularizer preventing source-domain overfitting. At an aggressive 90% sparsity (31.9M parameters, 23.3G FLOPs), the model maintains an average performance drop of only 1.30% across multiple benchmarks.

Ablation studies at 75% sparsity reveal that removing cross-domain distillation triggers catastrophic degradation on out-of-domain sets, driving In-the-Wild EER from 8.43% up to 25.84% and ASVspoof2019 LA from 3.03% to 10.14%. Removing movement-guided pruning moderately degrades performance on complex high-fidelity benchmarks like ASVspoof5 (17.75% to 19.95%) and FoR (10.51% to 13.56%), confirming its efficacy in preserving hard-to-detect forgery artifacts under high compression.

Systems / ConditionsSparsity#ParamsFLOPs19LA Eval EER%21DF EER%ASV5 EER%
XLSR-AASIST (Dense)0%317 M146.3 G0.143.0919.92
Finetune-Pruning [20]75%79.0 M43.7 G0.388.2326.13
Hybrid-Pruning [21]75%78.1 M50.3 G13.1930.5044.73
Ours75%79.2 M43.9 G0.222.7617.75
Ours90%31.9 M23.3 G0.303.7820.16

Limitations

The cross-domain distillation strategy is fundamentally bounded by the teacher model's error rate on unlabeled out-of-domain data, making further gains difficult when teacher confidence is low. The framework was exclusively validated on the XLS-R architecture and AASIST backend, leaving cross-architecture portability unverified. Furthermore, the study does not evaluate scaling to ultra-large speech SSL backbones like XLS-R-2B.

Why read this

Speech and ML engineers looking to deploy high-performance SSL-based audio deepfake detectors onto resource-constrained edge hardware will find this a definitive recipe. It provides concrete structural insights into how Transformer FFN, MHA, and CNN blocks behave under aggressive pruning when paired with multi-domain distillation.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

On-device audio deepfake detection, real-time voice cloning countermeasures for mobile operating systems, secure telephony fraud prevention, and edge-based speaker verification guardrails.

Institutions

Zhejiang University, Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Shanghai Institute for Advanced Study, China University of Petroleum (East China)

Funding / 經費: Shanghai Municipal Special Program for Basic Research on General AI Foundation Models, National Natural Science Foundation of China, Zhejiang Provincial Natural Science Foundation of China

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1766