All papers
Deepfakes & securityFull-paper digest

Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?

Xin Wang, Wanying Ge, Junichi Yamagishi

Code & resourcesgithub.com/nii-yamagishilab/AntiDeepfake

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.1 KB · Ready to paste

Preview copied content

TL;DR — The paper investigates applying Group Relative Policy Optimization (GRPO) instead of Supervised Fine-Tuning (SFT) to post-trained speech foundation models for binary deepfake detection, achieving better out-of-domain generalization while preserving in-domain performance.

Key contributions

  • Demonstrated that pure GRPO fine-tuning on SSL-based post-trained speech deepfake detectors improves or maintains out-of-domain generalization without hurting target-domain accuracy.
  • Showed that SFT followed by GRPO (hybrid training) underperforms compared to pure GRPO-only fine-tuning.
  • Identified via ablation studies that the negative reward component within GRPO's group-normalized advantage calculation is crucial for generalization improvements.
  • Provided Wasserstein-distance data drift analyses showing that GRPO prevents test distribution drift on out-of-domain evaluation sets compared to SFT.

Problem

Speech deepfake detectors built via supervised fine-tuning (SFT) on foundation models suffer from catastrophic forgetting, severely degrading on out-of-domain (OOD) test sets despite performing well in-domain. While large language models successfully leverage reinforcement learning (like GRPO) to align models and preserve pre-trained knowledge, speech detection has remained locked into SFT-only paradigms. This limits real-world deployability where acoustic conditions and deepfake generation algorithms are constantly evolving and unseen during training.

Method

The detection architecture combines a multi-lingual self-supervised learning (SSL) front-end (XLS-R-2B, MMS-1B, or MMS-300M) whose final-layer features are pooled via global average pooling, followed by a linear binary classifier outputting FAKE/REAL probabilities. The paper investigates fine-tuning this post-trained parameter set Θ using reinforcement learning via Group Relative Policy Optimization (GRPO). Unlike sequence-level LLM outputs, the detector outputs a single binary label, and the reward function is defined as a simple indicator function r(y_tilde, y) = delta(y_tilde == y). GRPO samples a group of G=64 outputs for each input, computes group-normalized advantages utilizing both positive and negative rewards, and avoids full RLHF overhead by utilizing simplified GRPO (GRPOs) where clipping and separate old-policy snapshots can be bypassed.

The training recipe uses a maximum of 10 epochs with validation early-stopping every 20k steps on the Deepfake-Eval-2024 (DFE24) training partition (~50 hours). For regularization, a Kullback-Leibler divergence penalty term with weight beta = 0.04 measures the distance against a frozen reference model (the post-trained checkpoint). Variants are tested by altering beta (beta=0 for no regularization, beta=1 for heavy regularization) and removing negative rewards to isolate the source of performance gains.

Experimental setup

Experiments use the DFE24 training set (~50 hours) for fine-tuning, and evaluate on DFE24 validation partitions (segmented into lengths of 4s, 10s, 13s, 30s, and 50s) as in-domain sets, plus ADD23, FoR, DV, and ItW datasets as out-of-domain test sets using Equal Error Rate (EER) as the primary metric. The base models evaluated include XLS-R-2B, MMS-1B, and MMS-300M pre- and post-trained via the AntiDeepfake project. Training is executed on an H100 GPU (and TSUBAME4.0 supercomputer), running each experiment across three random rounds and averaging the resulting EERs.

Results

Pure GRPO on the post-trained XLS-R-2B model achieves a 9.93% average in-domain EER (comparable to SFT's 10.26%), but drastically outperforms SFT on out-of-domain test sets, dropping EER on the In-the-Wild (ItW) set from 6.35% (SFT) down to 2.19% (GRPO). SFT-then-GRPO hybrid fine-tuning fails to beat pure GRPO, yielding higher OOD EERs such as 5.89% on ItW and 7.04% on DV. Ablations indicate that removing negative rewards degrades performance across all datasets, while setting the KL-penalty weight too high (beta=1.0) causes severe underfitting, raising in-domain EER to 15.83%. GRPO fails to yield these generalization benefits if applied directly to raw pre-trained models without intermediate post-training.

System / ConditionDFE24 (4s)DFE24 (ave.)ADD23FoRDVItW (ave.)
Post-trained only (XLS-R-2B)27.7322.644.672.612.232.69
SFT-only (XLS-R-2B)12.1710.266.093.928.756.35
GRPO-only (XLS-R-2B, beta=0.04)11.069.935.340.472.762.69
SFT -> GRPO (XLS-R-2B)11.539.815.782.757.045.37
GRPO w/o negative reward13.3011.416.690.833.093.44
GRPO (beta = 1.0)19.3715.834.021.341.662.04

Limitations

The study is restricted to binary classification architectures using fixed-dimensional global average pooling over SSL front-ends, omitting sequence-to-sequence frame-level localization tasks. The fine-tuning target domain is limited to DFE24 (~50 hours), and broader languages or specialized vocoding spoofing types beyond the tested datasets remain unverified. Additionally, computing multiple rollouts per sample (G=64) increases training time relative to standard supervised cross-entropy optimization.

Why read this

Speech and ML engineers building robust speech deepfake detectors should read this to learn how to transition from standard supervised fine-tuning to reinforcement learning (GRPO) to mitigate catastrophic forgetting and dramatically improve out-of-domain generalization.

Code

Applications

Robust speech anti-spoofing systems, telephony fraud detection pipelines, and automated media verification tools operating across diverse, unseen acoustic environments.

Institutions

National Institute of Informatics

Funding / 經費: Japan Science and Technology Agency, New Energy and Industrial Technology Development Organization

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-589