All papers
Audio understandingFull-paper digest

The silence of the weights: a structural pruning strategy for Attention-based audio signal architectures with second-order metrics

Andrea Diecidue, Carlo Alberto Barbano, Piero Fraternali, Mathieu Fontaine, Enzo Tartaglione

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.2 KB · Ready to paste

Preview copied content

TL;DR — This paper introduces a structured per-head channel pruning strategy combined with Fisher information scoring for attention-based audio architectures, preserving performance within 1% even at 50% sparsity.

Key contributions

  • Proposes a per-head (PH) channel-wise structured pruning scheme that independently selects channels to prune per head under matrix dimension constraints (Wq/WkW_q/W_k and Wv/WoW_v/W_o).
  • Adopts a linear-cost Fisher information (FI) second-order scoring metric to overcome parameter magnitude scale biases across layers.
  • Compares global (G) and local (L) thresholding strategies across diverse tasks using Audio Spectrogram Transformer (AST) and Whisper models.
  • Demonstrates that 50% parameter removal in attention blocks maintains competitive accuracy (e.g., 97.71% on SpeechCommands) compared to whole head-wise pruning.

Problem

Standard transformer models used in machine listening scale up to billions of parameters, resulting in high energy, memory, and inference latency requirements. Traditional structured pruning methods focus heavily on coarse whole-head pruning (EH) or token dropping, while unstructured approaches like PARP fail to deliver actual wall-clock speedups. Furthermore, simple magnitude-based pruning metrics suffer from cross-layer scale disparities, disproportionately damaging earlier layers unless carefully managed. Addressing fine-grained channel pruning within attention blocks for audio transformers remains largely underexplored.

Method

The proposed method focuses on the four self-attention weight matrices: Wq∈Rdq×dW_q \in \mathbb{R}^{d_q \times d}, Wk∈Rdq×dW_k \in \mathbb{R}^{d_q \times d}, Wv∈Rdv×dW_v \in \mathbb{R}^{d_v \times d}, and Wo∈Rd×dvW_o \in \mathbb{R}^{d \times d_v}. Pruning maintains two strict dimensional constraints: WqW_q and WkW_k must share an equal output dimension, and WvW_v's output dimension must match WoW_o's input dimension. Unlike standard Entire Head (EH) pruning that removes entire attention heads, the Per-Head (PH) channel-wise approach independently allocates a sparsity budget per head to eliminate redundant channels while preserving crucial subspaces.

To score parameters without being misled by magnitude variations across layers, the method computes the Fisher Information (FI) using a log-likelihood loss over a sample dataset XX. Unlike the full Hessian matrix which scales quadratically, Fisher information provides second-order sensitivity information in linear time. For thresholding, the framework investigates global (G) strategies—which pool all channels across layers to allocate a flexible sparsity per layer—and local (L) strategies, which enforce uniform sparsity across all layers.

During training, AST models are fine-tuned iteratively using LoRA for 3 epochs per sparsity step with AdamW (lr=10−4lr=10^{-4}), while Whisper medium models are fine-tuned across a 33k-hour multilingual audio corpus (LibriSpeech, MLS, CommonVoice, VoxPopuli, FLEURS, CoVoST) using SGD (lr=10−4lr=10^{-4}).

Experimental setup

Evaluated on AudioSet (balanced subset) and SpeechCommands v2 for AST classification, alongside Whisper medium evaluated on LibriSpeech (English), CommonVoice (Italian, French), and CoVoST (German-to-English translation). Pruning is executed iteratively across 10 steps, removing 10% of attention parameters per iteration up to 50-60% sparsity. Baselines include Entire Head (EH) pruning, magnitude-based (MAG) scoring, and local vs. global thresholding variants.

Results

At 60% sparsity on SpeechCommands, the proposed Fisher-guided per-head approach achieved 97.71% accuracy (vs. 97.51% for head-wise Fisher), and 30.86 mAP on AudioSet (vs. 31.10 mAP for head-wise Fisher). Magnitude-based pruning required local thresholding (97.49% on SpeechCommands) to avoid collapse, whereas Fisher information excelled under global thresholding due to its scale invariance. In terms of latency, head-wise pruning yielded slightly faster inference (1-2 ms faster than per-head) because it completely removes scaled dot-product operations from the computation graph. Machine translation tasks (CoVoST DE-EN) showed larger performance drops due to small domain-specific fine-tuning sets (~1000 samples).

System / ConditionSparsitySpeechCommands Acc (%)AudioSet mAPLibriSpeech WER
Baseline (Unpruned)0%~98.0~32.0Baseline
PHGFI60%97.71
EHGFI60%97.51
PHLMAG60%97.49
EHGMAG60%96.54

Limitations

The study is restricted to attention blocks, omitting feed-forward networks (FFNs) from the pruning loop. Fine-tuning datasets for low-resource translation tasks were relatively small (~1000-1500 samples), exacerbating performance degradation at higher sparsities. Per-head channel pruning yields lesser raw hardware speedups compared to full head-removal because inner dimension contractions do not entirely eliminate structural operators.

Why read this

Speech and machine learning engineers seeking to compress large audio transformer architectures (like Whisper or AST) without retraining from scratch should read this to understand how second-order Fisher metrics and per-head channel pruning outperform naive magnitude scoring.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Deploying resource-constrained automatic speech recognition (ASR) and audio classification models on edge or on-device hardware.

Institutions

Politecnico di Milano, University of Turin, Telecom Paris, Institut Polytechnique de Paris

Funding / 經費: French National Research Agency, Hi! PARIS, ANR/France 2030 program, Italian Ministry of University and Research, European Union

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2026