All papers
Speaker recognitionFull-paper digest

Continuous 2D Spectral—Temporal Transformer for Speaker Verification

Seongwook Ham, Thien-Phuc Doan

Code & resourcesgithub.com/roadroller0501/C2D-ST

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

11.9 KB · Ready to paste

Preview copied content

TL;DR — The Continuous 2D Spectral–Temporal Transformer (C2D-ST) preserves the explicit 2D time-frequency grid structure throughout its feature extraction backbone rather than collapsing the frequency axis prematurely, achieving a competitive 0.507 average EER on VoxCeleb with only 6.9M parameters.

Key contributions

  • Proposes a continuous spectral-temporal attention backbone that preserves the 2D grid format across five transformer stages before applying temporal-only aggregation.
  • Introduces axial attention operating along both the temporal and frequency axes directly within the spectral-temporal domain for global dependency modeling.
  • Combines neighborhood attention with value-side relative positional encoding (RPE) and relative position bias (RPB) for robust local spatial modeling.
  • Demonstrates superior parameter efficiency compared to larger hybrid baselines like ReDimNet-B6 and ECAPA2 while reducing average EER.

Problem

Recent hybrid speaker verification systems frequently collapse the spectral axis into a 1D temporal format early or intermittently during stage-wise processing (e.g., ReDimNet), destroying the explicit 2D spectral-temporal grid structure prior to global dependency modeling. This premature flattening limits the network's capacity to learn comprehensive 2D contextual patterns and degrades parameter efficiency. Addressing this gap requires examining whether maintaining an uninterrupted 2D spectral-temporal representation during global dependency learning yields more discriminative and compact speaker embeddings.

Method

The C2D-ST architecture accepts 80-dimensional log Mel filterbank features (B×C×F×TB \times C \times F \times T) and processes them through a backbone comprising 5 spectral-temporal attention stages totaling 16 transformer layers. Each stage incorporates Neighborhood Attention (NA) for local 2D spatial context using relative position bias (RPB) and value-side relative positional encoding (RPE), followed by a stage-final Axial Attention layer that decomposes global dependencies into separate frequency and temporal branches.

Unlike architectures with frequent format interleaving, C2D-ST maintains the 2D grid throughout the backbone. The outputs of multiple backbone stages are aggregated via learnable scalar weights, projected to a 1D channel format, and passed to a final temporal modeling stage consisting of 8 temporal layers (NA and temporal Global Attention). Finally, Attentive Statistics Pooling (ASP) extracts the speaker embedding.

The model is pretrained from scratch on VoxCeleb2 using random 3-second segments for 20 epochs with AdamW (initial lr=0.008, weight decay=0.05) under SphereFace2 (C-type) loss with margin 0.2, utilizing speed perturbation, MUSAN noise, and simulated RIRs. Large-margin finetuning (LMFT) is subsequently performed for 2 epochs on 6-second segments with lr=0.003 and margin 0.3 without augmentations.

Experimental setup

Trained on the VoxCeleb2 development set and evaluated on VoxCeleb1 (VoxCeleb1-O, VoxCeleb1-E, and VoxCeleb1-H test protocols). Evaluated using Equal Error Rate (EER) and minimum Detection Cost Function (minDCF), alongside adaptive score normalization (AS-Norm with 500 cohort speakers) and quality-measure scoring (QMF). Implemented using an adaptation of ESPnet-SPK.

Results

C2D-ST achieves an average EER of 0.507 and minDCF of 0.051 on VoxCeleb1 using 6.9M parameters, outperforming ReDimNet-B6 (15.0M params, 0.633 average EER) and ECAPA2 (27.1M params, 0.617 average EER). Ablation studies reveal that removing axial attention and relying solely on neighborhood attention spikes the average EER to 0.693, whereas restricting axial attention to time-only yields 0.533. Replacing the proposed 2D neighborhood attention blocks with ConvNeXt convolutional blocks increases average EER to 0.580.

ModelParamsLMFTQMFVox1-OVox1-EVox1-HAvgEERAvgminDCF
ReDimNet-B6 [7]15.0M✓✗0.37 / 0.0300.53 / 0.0511.00 / 0.0970.6330.059
ECAPA2 [3]27.1M✓✓0.34 / 0.0290.52 / 0.0580.99 / 0.0980.6170.062
C2D-ST6.9M✗✗0.37 / 0.0390.55 / 0.0540.96 / 0.0930.6270.062
+ LMFT6.9M✓✗0.31 / 0.0340.44 / 0.0450.80 / 0.0770.5170.052
+ LMFT + QMF6.9M✓✓0.29 / 0.0340.44 / 0.0440.79 / 0.0760.5070.051

Limitations

Evaluated exclusively on the VoxCeleb benchmark dataset, leaving generalization to heavily noisy, reverberant real-world telephony, or cross-lingual scenarios unexplored. The computational overhead of continuous 2D attention grids on extremely long audio streams or edge devices is not measured.

Why read this

Researchers building high-efficiency speaker verification models should read this to understand how preserving continuous 2D spectral-temporal grids with axial attention outperforms early frequency-collapse strategies.

Code

Applications

Speaker verification, speaker recognition, and voice biometrics systems requiring high accuracy under constrained parameter budgets.

Institutions

Soongsil University

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-963