---
id: yin26_interspeech
title: "The First Environmental Sound Deepfake Detection Challenge: Benchmarking
  Robustness, Evaluation, and Insights"
authors:
  - Han Yin
  - Yang Xiao
  - Rohan Kumar Das
  - Jisheng Bai
  - Ting Dang
year: 2026
doi: 10.21437/Interspeech.2026-1599
isca_url: https://www.isca-archive.org/interspeech_2026/yin26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/yin26_interspeech.pdf
session: Evaluation, Benchmarking, and Reliability of Audio Systems
topics:
  - audio-deepfake
  - self-supervised
  - evaluation
category: deepfake-security
labels:
  - low-resource
  - self-supervised
  - dataset-or-benchmark-release
  - robustness-noise
institutions:
  - KAIST
  - University of Melbourne
  - Fortemedia Singapore
  - Xi'an University of Posts & Telecommunications
  - Xi'an Lianfeng Acoustic Technologies Co., Ltd
code:
  url: https://github.com/apple-yinhan/ESDD-Review
  stars: 1
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: yin26_interspeech
  category: deepfake-security
  labels:
    - low-resource
    - self-supervised
    - dataset-or-benchmark-release
    - robustness-noise
  institutions:
    - KAIST
    - University of Melbourne
    - Fortemedia Singapore
    - Xi'an University of Posts & Telecommunications
    - Xi'an Lianfeng Acoustic Technologies Co., Ltd
  code: https://github.com/apple-yinhan/ESDD-Review
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-1599
  pdf: https://www.isca-archive.org/interspeech_2026/yin26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/yin26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/yin26_interspeech/markdown.md
---

# The First Environmental Sound Deepfake Detection Challenge: Benchmarking Robustness, Evaluation, and Insights

*Han Yin, Yang Xiao, Rohan Kumar Das, Jisheng Bai, Ting Dang*

[PDF](https://www.isca-archive.org/interspeech_2026/yin26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/yin26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-1599)

**Category:** `deepfake-security` · **Labels:** `low-resource`, `self-supervised`, `dataset-or-benchmark-release`, `robustness-noise`

**TL;DR** — This paper presents the benchmark results and analysis of the first Environmental Sound Deepfake Detection (ESDD) challenge, featuring 97 teams and 1,748 submissions across unseen generator and black-box low-resource tracks. Top-performing ensembles achieved Equal Error Rates (EER) as low as 0.30% by leveraging advanced self-supervised audio representations and domain-targeted data augmentations.

## Key contributions

- Formulated the first systematic benchmark and dataset (EnvSDD) for environmental sound deepfake detection covering monophonic and polyphonic soundscapes.
- Designed two rigorous challenge tracks: Track 1 for unseen text-to-audio and audio-to-audio generators, and Track 2 for black-box video-to-audio generators under severe low-resource data constraints (1% training set size).
- Analyzed top-performing architectures, demonstrating that pre-trained front-ends (EAT, SSLAM, BEATs) combined with advanced classification backends (BiCrossMamba, multi-branch AASIST) dramatically outperform baseline models.
- Highlighted the vulnerability of baseline detectors to advanced flow-matching generators like TangoFlux (reaching >20% EER) and established robust ensemble strategies to mitigate this degradation.

## Problem

Detecting deepfakes has historically focused almost exclusively on human speech and singing voices, leaving environmental sound deepfake detection (ESDD) underexplored. Environmental sounds encompass diverse, complex monophonic and polyphonic acoustic scenes with overlapping events, making traditional pitch or phonetic anomaly cues unreliable. Furthermore, existing detectors degrade catastrophically when evaluated on unseen text-to-audio (TTA), audio-to-audio (ATA), and video-to-audio (VTA) generators, necessitating robust, generator-agnostic architectures that function under data scarcity.

## Method

Top-tier systems generally adopted a two-stage paradigm: extracting high-level acoustic representations using advanced self-supervised learning (SSL) models (such as EAT, SSLAM, or BEATs) followed by specialized temporal-spectral classifiers. Architecturally, winning solutions utilized double AASIST branches processing frequency and channel splits, bidirectional selective state space models (BiCrossMamba) for spectro-temporal modeling, or lightweight multi-head factorized attention (MHFA) modules. Training strategies frequently incorporated class-weighted objectives to address severe real-versus-fake data imbalances, LoRA fine-tuning, domain adversarial training, and ArcFace loss. To survive open-world distribution shifts, participants engineered targeted augmentations including MP3 compression, loudness normalization, Cut and Mix, and mixing in external real data from AudioCaps, ultimately relying on multi-system ensembles to capture complementary generator-agnostic spoofing artifacts.

## Experimental setup

Track 1 used the EnvSDD dataset (27,811 real and 111,244 fake training clips generated by AudioLDM, AudioLDM 2, and AudioGen; tested on 1,000 real and 3,000 fake clips from unseen AudioLCM and TangoFlux). Track 2 introduced a low-resource black-box scenario with only 270 real and 1,083 fake training clips (1% of the development pool) paired against 1,994 real and 7,980 test clips synthesized via video-to-audio models FoleyCrafter and DiffFoley. Evaluation was performed using Equal Error Rate (EER) against baseline models AASIST and BEATs+AASIST.

## Results

The best ensemble system in Track 1 (AHU team, S01) achieved an exceptional EER of 0.30%, compared to 13.20% for the BEATs+AASIST baseline and 15.02% for vanilla AASIST. TangoFlux proved to be the most formidable unseen TTA generator, driving baseline EERs up to 19.00%-20.10%, whereas top ensembles like S01 suppressed this error down to 0.30% via specialized data augmentation and multi-classifier fusion. In Track 2's low-resource black-box setting, top-ranking systems (DFKI M01 and AHU M02) matched this performance with a 0.25% EER, cleanly outperforming the baseline's 12.48% EER.

| System / Condition | Track 1 EER (%) | Track 2 EER (%) |
|---|---|---|
| Baseline 2 (AASIST) | 15.02 | 15.40 |
| Baseline 1 (BEATs+AASIST) | 13.20 | 12.48 |
| CAU (BEAT2AASIST Ensemble) | 1.60 | 0.35 |
| DFKI (EAT-L+BiCrossMamba) | 0.80 | 0.25 |
| AHU (EAT+AASIST Ensemble) | 0.30 | 0.25 |

## Limitations

The evaluation relies on clip-level binary classification (real vs. fake), which fails to handle localized manipulations within polyphonic mixtures where only specific background or foreground components are forged. Track 2 simulated low-resource settings using only a 1% data split from two specific video-to-audio generators, which may not capture the full diversity of real-world black-box editing tools. Furthermore, dependence on large external SSL pre-trained encoders creates heavy computational overhead, limiting real-time on-device edge deployment.

## Why read this

Researchers and engineers working on audio forensics and anti-spoofing should read this to understand how pre-trained SSL features, state space models (Mamba), and aggressive augmentation pipelines can be adapted to secure environmental soundscapes against generative audio models.

## Code

- https://github.com/apple-yinhan/ESDD-Review

## Applications

Public safety verification, surveillance audit systems, digital media forensics, and automated moderation for user-uploaded audio-visual platforms.

## Institutions / 機構

KAIST, University of Melbourne, Fortemedia Singapore, Xi'an University of Posts & Telecommunications, Xi'an Lianfeng Acoustic Technologies Co., Ltd

## Related

- [FakeSound2: A Benchmark for Explainable, Traceable, and Generalizable Deepfake Sound Detection](xie26_interspeech.md) — shared data / evaluation · relatedness 2.4/3
- [GradHarmony: A Gradient Alignment and Magnitude Normalization Strategy for Audio Deepfake Detection](kim26r_interspeech.md) — same problem · relatedness 2.2/3
- [Diffusion Reconstruction towards Generalizable Audio Deepfake Detection](cheng26_interspeech.md) — same problem · relatedness 2.2/3
- [DeepFense: A Unified, Modular, and Extensible Framework for Robust Audio Deepfake Detection](kheir26_interspeech.md) — shared technique · relatedness 2.2/3
- [ADD-DINO: A Two-Stage Self-Distillation Framework for Audio Deepfake Detection](sun26g_interspeech.md) — same problem · relatedness 2.1/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
