All papers
Deepfakes & securityFull-paper digest

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus

Kirill Borodin, Vasiliy Kudryavtsev, Maxim Maslov, Mikhail Gorodnichev, Grach Mkrtchian

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.9 KB · Ready to paste

Preview copied content

TL;DR — The paper introduces LRLspoof, a 2,732-hour multilingual corpus spanning 66 languages (45 low-resource) and 24 TTS systems, revealing through zero-shot threshold transfer that language mismatch is an independent source of domain shift causing severe performance disparities in spoof detection.

Key contributions

  • Introduces LRLspoof, comprising 2,732.17 hours of audio across 66 languages generated by 24 open-source TTS systems, featuring 45 low-resource languages.
  • Establishes a unified zero-shot evaluation protocol using threshold transfer calibrated on pooled external benchmarks (ASVspoof5, ASVspoof2021 LA/DF, In-the-Wild, DFADD, ADD2022) to report Spoof Rejection Rate (SRR) without requiring target-domain bona fide speech.
  • Demonstrates that language identity acts as an independent driver of domain shift, showing performance variations of tens of percentage points even when the TTS generator and countermeasure are held fixed.
  • Provides comprehensive evaluations and aggregated metrics (mean and pooled SRR) across 11 publicly available spoofing countermeasures, highlighting vulnerability in low-resource settings.

Problem

Current speech anti-spoofing benchmarks and countermeasures (CMs) are heavily dominated by a small set of high-resource languages, causing models to rely on language- or phonotactics-correlated artifacts rather than robust spoofing cues. Prior generalization studies focus mostly on unseen synthesizers, vocoders, or codecs, leaving language mismatch largely underexplored as a latent bias factor. This matters operationally because modern voice cloning and TTS tools are increasingly multilingual, yet deployed detectors fail unpredictably when facing linguistically heterogeneous audio streams.

Method

The LRLspoof corpus was constructed using only synthetically generated speech produced by 24 upstream open-source TTS implementations executed with default configurations and public pretrained weights, requiring no codebase modifications. The systems span three categories: classical non-neural parametric models (e.g., eSpeak NG, RHVoice, AhoTTS), neural supervised models with fixed voices (e.g., Silero, SpeechT5, FastPitch, Matcha-TTS, Parler-TTS, Piper, MeloTTS, MMS-TTS, Indic-TTS, TurkicTTS, IMS Toucan, QirimtatarTTS), and generative zero-shot voice cloning models (e.g., XTTS, XTTS2, OuteTTS, Chatterbox, F5-TTS, CosyVoice, Zonos, Fish-Speech, Kokoro). Text prompts were sourced from public multilingual text collections to preserve natural sentence structures and short utterance lengths suitable for evaluation.

Because LRLspoof contains no bona fide speech, the authors evaluate countermeasures using a zero-shot threshold transfer protocol. An Equal Error Rate (EER) operating threshold tau_EER is calibrated by pooling external benchmarks containing both genuine and fake speech (ASVspoof5, ASVspoof2021 LA/DF, In-the-Wild, DFADD, and ADD2022). This single fixed threshold is then applied directly to each language/synthesizer subset of LRLspoof to compute the Spoof Rejection Rate (SRR). This design deliberately avoids supervised per-language threshold tuning or introducing domain-confused bona fide target data, serving purely as a spoof-side cross-lingual stress test.

Experimental setup

Evaluates 11 publicly available spoofing countermeasures: aasist3, df_arena_1b, df_arena_500, res2tcn, rescapsguard, sls, ssl_aasist, tcm_add, nes2net, w2v2_1b, and w2v2_300. The evaluation dataset comprises 2,732.17 hours across 66 languages (45 low-resource, defined as having less than 100 hours of scripted speech in Common Voice 24.0) and 24 TTS systems. Metrics reported include Equal Error Rate (EER, %) on external calibration sets, and Spoof Rejection Rate (SRR, %) aggregated as mean SRR across languages (MSRR) and pooled SRR across utterances (PSRR) for both all languages and the low-resource subset.

Results

Overall evaluation shows extreme cross-lingual disparity across models under transferred EER thresholds. For example, aasist3 achieves high SRR on English (93.33%) and Chechen (99.86%), whereas w2v2_300 drops sharply to 62.27% on English and 30.78% on Polish. Aggregated metrics show that aasist3 achieves the highest overall performance (MSRR All: 90.40%, PSRR All: 91.40%, MSRR Low: 93.28%, PSRR Low: 94.28%), followed by w2v2_300 (MSRR All: 80.51%, PSRR All: 78.32%). Conversely, simpler or older architectures struggle significantly, with w2v2_1b achieving an MSRR of only 26.79% and res2tcn achieving 33.07%. Controlled experiments holding the TTS model constant reveal massive performance gaps due to language alone; under Parler-TTS, the English-Polish contrast yields a 94.46 pp gap for w2v2_300, and under Piper, the Danish-Welsh contrast yields a 99.51 pp gap for aasist3. Similar drops occur in low-resource pairings, such as a 75.38 pp gap between Hindi and Odia under Indic-TTS for rescapsguard.

ModelMSRR (All)PSRR (All)MSRR (Low)PSRR (Low)
aasist390.4091.4093.2894.28
df_arena_1b71.0770.5373.6274.89
df_arena_50063.4758.5067.3162.47
ssl_aasist45.6243.4850.0949.34
w2v2_30080.5178.3283.9081.29

Limitations

The evaluation is strictly spoof-only and does not measure False Rejection Rates (FRR) or target-domain EER because adding external bona fide speech would introduce domain confounds. The zero-shot protocol relies entirely on external calibration benchmarks, which may not adequately cover all deployment languages, recording channels, or acoustic environments. Furthermore, low-resource subsets depend on synthetic generation defaults which may not reflect human-spoken accents or diverse dialectal variations in those target regions.

Why read this

Speech security researchers and engineers building deployable anti-spoofing systems should read this paper to understand that cross-lingual transfer is a major blind spot in current countermeasures. It provides a rigorous framework and diagnostic dataset (LRLspoof) for stress-testing detectors against language-induced domain shifts rather than evaluating purely on high-resource benchmarks.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Auditing and improving the robustness of speech-driven security systems, speaker verification pipelines, and audio deepfake detectors against multilingual and low-resource spoofing attacks.

Institutions

Moscow Technical University of Communications and Informatics, BitmanagerAI

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-345