TL;DR — This paper presents the benchmark results and analysis of the first Environmental Sound Deepfake Detection (ESDD) challenge, featuring 97 teams and 1,748 submissions across unseen generator and black-box low-resource tracks. Top-performing ensembles achieved Equal Error Rates (EER) as low as 0.30% by leveraging advanced self-supervised audio representations and domain-targeted data augmentations.
Key contributions
- Formulated the first systematic benchmark and dataset (EnvSDD) for environmental sound deepfake detection covering monophonic and polyphonic soundscapes.
- Designed two rigorous challenge tracks: Track 1 for unseen text-to-audio and audio-to-audio generators, and Track 2 for black-box video-to-audio generators under severe low-resource data constraints (1% training set size).
- Analyzed top-performing architectures, demonstrating that pre-trained front-ends (EAT, SSLAM, BEATs) combined with advanced classification backends (BiCrossMamba, multi-branch AASIST) dramatically outperform baseline models.
- Highlighted the vulnerability of baseline detectors to advanced flow-matching generators like TangoFlux (reaching >20% EER) and established robust ensemble strategies to mitigate this degradation.
Problem
Detecting deepfakes has historically focused almost exclusively on human speech and singing voices, leaving environmental sound deepfake detection (ESDD) underexplored. Environmental sounds encompass diverse, complex monophonic and polyphonic acoustic scenes with overlapping events, making traditional pitch or phonetic anomaly cues unreliable. Furthermore, existing detectors degrade catastrophically when evaluated on unseen text-to-audio (TTA), audio-to-audio (ATA), and video-to-audio (VTA) generators, necessitating robust, generator-agnostic architectures that function under data scarcity.
Method
Top-tier systems generally adopted a two-stage paradigm: extracting high-level acoustic representations using advanced self-supervised learning (SSL) models (such as EAT, SSLAM, or BEATs) followed by specialized temporal-spectral classifiers. Architecturally, winning solutions utilized double AASIST branches processing frequency and channel splits, bidirectional selective state space models (BiCrossMamba) for spectro-temporal modeling, or lightweight multi-head factorized attention (MHFA) modules. Training strategies frequently incorporated class-weighted objectives to address severe real-versus-fake data imbalances, LoRA fine-tuning, domain adversarial training, and ArcFace loss. To survive open-world distribution shifts, participants engineered targeted augmentations including MP3 compression, loudness normalization, Cut and Mix, and mixing in external real data from AudioCaps, ultimately relying on multi-system ensembles to capture complementary generator-agnostic spoofing artifacts.
Experimental setup
Track 1 used the EnvSDD dataset (27,811 real and 111,244 fake training clips generated by AudioLDM, AudioLDM 2, and AudioGen; tested on 1,000 real and 3,000 fake clips from unseen AudioLCM and TangoFlux). Track 2 introduced a low-resource black-box scenario with only 270 real and 1,083 fake training clips (1% of the development pool) paired against 1,994 real and 7,980 test clips synthesized via video-to-audio models FoleyCrafter and DiffFoley. Evaluation was performed using Equal Error Rate (EER) against baseline models AASIST and BEATs+AASIST.
Results
The best ensemble system in Track 1 (AHU team, S01) achieved an exceptional EER of 0.30%, compared to 13.20% for the BEATs+AASIST baseline and 15.02% for vanilla AASIST. TangoFlux proved to be the most formidable unseen TTA generator, driving baseline EERs up to 19.00%-20.10%, whereas top ensembles like S01 suppressed this error down to 0.30% via specialized data augmentation and multi-classifier fusion. In Track 2's low-resource black-box setting, top-ranking systems (DFKI M01 and AHU M02) matched this performance with a 0.25% EER, cleanly outperforming the baseline's 12.48% EER.
| System / Condition | Track 1 EER (%) | Track 2 EER (%) |
|---|---|---|
| Baseline 2 (AASIST) | 15.02 | 15.40 |
| Baseline 1 (BEATs+AASIST) | 13.20 | 12.48 |
| CAU (BEAT2AASIST Ensemble) | 1.60 | 0.35 |
| DFKI (EAT-L+BiCrossMamba) | 0.80 | 0.25 |
| AHU (EAT+AASIST Ensemble) | 0.30 | 0.25 |
Limitations
The evaluation relies on clip-level binary classification (real vs. fake), which fails to handle localized manipulations within polyphonic mixtures where only specific background or foreground components are forged. Track 2 simulated low-resource settings using only a 1% data split from two specific video-to-audio generators, which may not capture the full diversity of real-world black-box editing tools. Furthermore, dependence on large external SSL pre-trained encoders creates heavy computational overhead, limiting real-time on-device edge deployment.
Why read this
Researchers and engineers working on audio forensics and anti-spoofing should read this to understand how pre-trained SSL features, state space models (Mamba), and aggressive augmentation pipelines can be adapted to secure environmental soundscapes against generative audio models.
Code
Applications
Public safety verification, surveillance audit systems, digital media forensics, and automated moderation for user-uploaded audio-visual platforms.
Institutions
KAIST, University of Melbourne, Fortemedia Singapore, Xi'an University of Posts & Telecommunications, Xi'an Lianfeng Acoustic Technologies Co., Ltd
Related
- FakeSound2: A Benchmark for Explainable, Traceable, and Generalizable Deepfake Sound Detection — shared data / evaluation · relatedness 2.4/3
- GradHarmony: A Gradient Alignment and Magnitude Normalization Strategy for Audio Deepfake Detection — same problem · relatedness 2.2/3
- Diffusion Reconstruction towards Generalizable Audio Deepfake Detection — same problem · relatedness 2.2/3
- DeepFense: A Unified, Modular, and Extensible Framework for Robust Audio Deepfake Detection — shared technique · relatedness 2.2/3
- ADD-DINO: A Two-Stage Self-Distillation Framework for Audio Deepfake Detection — same problem · relatedness 2.1/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1599