All papers
Speaker recognitionFull-paper digest

Soft-Gating Score-Level Fusion for Spoofing-Aware Speaker Verification

Seongkyu Han, Yowon Lee, Thien-Phuc Doan, Thien An Nguyen, Souhwan Jung

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.0 KB · Ready to paste

Preview copied content

TL;DR — A training-free soft-gating score-level fusion framework dynamically scales ASV and CM subsystem contributions using margins from development EER thresholds, achieving up to a 90% relative a-DCF improvement over static baselines.

Key contributions

  • Proposes a novel dynamic soft-gating score-level fusion method for Spoofing-Aware Speaker Verification (SASV) that relies on trial-wise subsystem confidence.
  • Eliminates the need for learnable parameters or additional joint training, allowing direct plug-and-play deployment on existing pre-trained SASV pipelines.
  • Evaluates three specific gating variants (CM Gating, ASV Gating, and Double Gating) across diverse benchmark conditions.
  • Provides diagnostic analysis of performance failure modes tied to extreme EER threshold distributions (near 0 or 1).

Problem

Combining Automatic Speaker Verification (ASV) and Countermeasure (CM) subsystems is challenging because they optimize for different objectives: speaker identity vs. spoof detection. Conventional static fusion schemes apply fixed weights across all trials, making them unable to adapt to trial-specific attack types or shifts in score distributions. More recent score-aware gated frameworks like ATMM-SAGA require complex joint training with alternating optimization, limiting their practical deployment.

Method

The method takes raw ASV and CM similarity scores, normalizes them, and scales them using confidence margins derived from development-set EER thresholds. Let sasvs_{\text{asv}} and scms_{\text{cm}} be the normalized subsystem scores, and τasv,τcm\tau_{\text{asv}}, \tau_{\text{cm}} be their respective EER thresholds. The confidence margins are computed as δasv=sasv−τasv\delta_{\text{asv}} = s_{\text{asv}} - \tau_{\text{asv}} and δcm=scm−τcm\delta_{\text{cm}} = s_{\text{cm}} - \tau_{\text{cm}}.

Three gating configurations are explored: CM Gating (S=scmδcm+sasv(1−∣δcm∣)S = s_{\text{cm}} \delta_{\text{cm}} + s_{\text{asv}}(1 - |\delta_{\text{cm}}|)), ASV Gating (S=scm(1−∣δasv∣)+sasvδasvS = s_{\text{cm}}(1 - |\delta_{\text{asv}}|) + s_{\text{asv}} \delta_{\text{asv}}), and Double Gating (S=scmδcm+sasvδasvS = s_{\text{cm}} \delta_{\text{cm}} + s_{\text{asv}} \delta_{\text{asv}}). These formulations scale subsystem contribution dynamically based on whether the score falls far from the threshold (high confidence) or close to it (uncertainty), giving strong influence to reliable outputs while suppressing ambiguous predictions.

At inference time, the method requires only simple arithmetic operations on the subsystem outputs using pre-calculated thresholds from the development set, requiring no backpropagation or parameter updates.

Experimental setup

Experiments are conducted on ASVspoof 2019 LA (LA19) and ASVspoof5 (Track 2 closed condition) datasets using their official evaluation protocols. Four ASV-CM system combinations are built using two ASV backbones (ECAPA-TDNN and ReDimNet, trained on VoxCeleb2) and two CM backbones (AASIST and Conformer-TCM). Performance is evaluated using SV-EER, SPF-EER, SASV-EER, and a-DCF.

Results

On the LA19 dataset, the proposed methods reduce a-DCF by approximately 90% on average compared to baseline fusion methods, with Double Gating achieving the best a-DCF across most configurations. On ASVspoof5, the dynamic gating methods also predominantly outperform Baseline 1 (simple sum) and Baseline 2 (DNN back-end embedding fusion). However, performance degrades severely when models produce extreme EER thresholds close to 0 or 1. For instance, in the ECAPA+TCM configuration where TCM's threshold is near 0, CM Gating yields an inflated SV-EER due to unconstrained amplification of the CM score on bonafide trials.

ModelMethodLA19 SASV-EERLA19 a-DCFASVspoof5 SASV-EERASVspoof5 a-DCF
ECAPA + AASISTBaseline 119.140.173835.030.6219
ECAPA + AASISTCM Gating0.840.017816.010.4883
ECAPA + TCMBaseline 18.120.150427.360.3065
ECAPA + TCMDouble Gating0.760.015129.270.8942
Redim + AASISTDouble Gating0.450.010713.680.4280
Redim + TCMDouble Gating0.390.008025.520.8113

Limitations

The framework's effectiveness is sensitive to the positioning of the development-set EER threshold; extreme thresholds near 0 or 1 break the margin scaling logic and degrade performance. The evaluation is restricted to closed-condition LA19 and ASVspoof5 benchmarks, leaving cross-dataset generalization under severe acoustic or codec mismatch unverified.

Why read this

Researchers and engineers looking for an immediate, plug-and-play alternative to static score-fusion or complex joint-training pipelines in SASV will find this an effective, lightweight solution.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Secure speaker verification systems, voice-biometric banking, and mobile authentication pipelines requiring defense against synthetic speech spoofing.

Institutions

Soongsil University

Funding / 經費: Institute of Information & Communications Technology Planning & Evaluation, Information Technology Research Center, Ministry of Science and ICT, National Research Foundation of Korea

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2294