---
id: sedaghati26_interspeech
title: "VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection
  under Perturbations"
authors:
  - Farnaz Sedaghati
  - Yuxi Wang
  - Zicheng Weng
  - Wei Rao
year: 2026
doi: 10.21437/Interspeech.2026-1771
isca_url: https://www.isca-archive.org/interspeech_2026/sedaghati26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/sedaghati26_interspeech.pdf
session: Audio Watermarking and Source Verification
topics:
  - audio-deepfake
  - evaluation
  - self-supervised
category: deepfake-security
labels:
  - multilingual
  - dataset-or-benchmark-release
  - robustness-noise
institutions:
  - University of Tehran
  - Nanyang Technological University
code:
  url: https://github.com/wailywang/VoxWatermark
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: sedaghati26_interspeech
  category: deepfake-security
  labels:
    - multilingual
    - dataset-or-benchmark-release
    - robustness-noise
  institutions:
    - University of Tehran
    - Nanyang Technological University
  code: https://github.com/wailywang/VoxWatermark
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-1771
  pdf: https://www.isca-archive.org/interspeech_2026/sedaghati26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/sedaghati26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/sedaghati26_interspeech/markdown.md
---

# VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection under Perturbations

*Farnaz Sedaghati, Yuxi Wang, Zicheng Weng, Wei Rao*

[PDF](https://www.isca-archive.org/interspeech_2026/sedaghati26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/sedaghati26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-1771)

**Category:** `deepfake-security` · **Labels:** `multilingual`, `dataset-or-benchmark-release`, `robustness-noise`

**TL;DR** — VoxWatermark is a massive audio watermarking benchmark spanning 126,513 hours across 25 languages, designed to evaluate detectors under rigorous no-box, black-box, and white-box perturbations. The authors also propose AudioWMD, a stochastic query-stability meta-detector that outperforms single-query baselines under white-box and clean OOD conditions.

## Key contributions

- Constructed VoxWatermark, covering 126,513.89 hours of audio, 25 languages, 91,090K samples, and 10 distinct watermarking methods (4 neural, 6 traditional).
- Introduced a deployment-oriented perturbation protocol encompassing no-box (17 signal distortions/codecs), black-box (HSJA, Square), and white-box attacks.
- Proposed AudioWMD, a two-stage detector using stochastic query transformations (K=8) and a 5-dimensional meta-feature logistic regression classifier.
- Open-sourced the full dataset and codebase to establish a standardized framework for audio provenance and watermark detection.

## Problem

As text-to-speech models achieve human-like realism, malicious impersonation and misinformation risks have escalated, making robust audio watermarking and detection critical. However, prior benchmarks (like AudioMarkBench and RAW-Bench) focus narrowly on watermark payload recovery or perceptual quality rather than systematic, detector-level evaluation under realistic distribution shifts and unknown embedding methods. Furthermore, existing detectors fail catastrophically when confronted with out-of-domain languages, unseen watermarking algorithms, and adversarial perturbations.

## Method

AudioWMD frames watermark detection as a binary classification task executed in a two-stage pipeline. In Stage I, a base detector processes 16 kHz log-mel spectrograms using a ConvNeXtV2-style backbone trained with binary cross-entropy loss on 61,200 clean and watermarked samples (covering 6 seen watermarking schemes without augmentations). In Stage II (query-statistics meta detection), each input clip undergoes K=8 stochastic transformations (including the original clip plus time-stretching, pitch-shifting, additive noise, gain scaling, and time-masking). The base detector generates a confidence score for each transformed variant.

From these K scores, AudioWMD extracts a compact 5-dimensional statistical feature vector consisting of the mean score, standard deviation, score range, positive occupancy ratio, and decision-flip rate relative to the original query. This vector is fed into a logistic-regression meta-classifier to output the final watermark probability. This design explicitly evaluates score consistency and stability under distribution shifts rather than relying on a single static forward pass.

## Experimental setup

The benchmark uses 60,000 unwatermarked clips (83.4 hours across LibriSpeech, Common Voice 25-language set, VCTK, and AISHELL-1) scaled up via 10 watermarking schemes to 126,513 hours total. Training uses 6 seen schemes (LSB, QIM, DSSS, AudioSeal, Timbre, Phase Coding); evaluation uses 2 OOD test sets featuring 3 unseen schemes (Patchwork, Echo, WavMark) with cross-lingual (non-EN/ZH Common Voice) and cross-accent (VCTK) splits. Models are compared against reproduced WMD using AUROC, Accuracy, Precision, Recall, and F1.

## Results

On the validation set, AudioWMD achieves an AUROC of 88.3% and Accuracy of 84.0%, outperforming the reproduced WMD baseline (72.0% AUROC, 67.0% Accuracy). On OOD Test Set 1, AudioWMD achieves an AUROC of 63.8% versus WMD's 57.1%. Under white-box attacks, AudioWMD shows massive resilience gains, reaching 77.15% AUROC on Test Set 1 compared to WMD's 48.63%. 

However, AudioWMD does not win everywhere: it struggles significantly with black-box gradient approximations like HSJA spec on Test Set 1, where its True Positive Rate drops to 3.91% compared to WMD's 96.09%. Under no-box background noise, codecs, and temporal shifts, both models experience severe performance degradation toward random chance levels.

| System | Dataset | AUROC (%) | Accuracy (%) | White-Box AUROC (%) |
|---|---|---|---|---|
| AudioWMD | Validation | 88.3 | 84.0 | - |
| WMD [27] | Validation | 72.0 | 67.0 | - |
| AudioWMD | Test Set 1 | 63.8 | 53.0 | 77.15 |
| WMD [27] | Test Set 1 | 57.1 | 55.0 | 48.63 |
| AudioWMD | Test Set 2 | 63.2 | 58.0 | 70.02 |
| WMD [27] | Test Set 2 | 57.9 | 56.0 | 41.18 |

## Limitations

Performance under no-box environmental, codec, and temporal perturbations still drops drastically near chance level for both detectors, highlighting a vulnerability to real-world transmission channels. The meta-detector exhibits a blind spot against specific black-box score-based and gradient-estimated attacks (e.g., HSJA spectrogram variants), proving that query-stability alone does not defend against all adversarial strategies. The evaluation is currently restricted to 16 kHz mono audio and fixed-length 5-second segments.

## Why read this

Read this paper if you build synthetic speech attribution systems or AIGC provenance filters and need a rigorous benchmark to test your detector's out-of-domain and adversarial robustness. You will take away a comprehensive multi-method watermarking dataset and architectural insights into why query-stability meta-features help against white-box gradients but falter against certain black-box attacks.

## Code

- https://github.com/wailywang/VoxWatermark

## Applications

AI-generated speech detection, media source attribution, content provenance verification, and combating audio deepfakes in open-world transmission environments.

## Institutions / 機構

University of Tehran, Nanyang Technological University

## Related

- [Learning to Evade: Adaptive Attacks on Audio Watermarking](ding26b_interspeech.md) — same problem · relatedness 2.3/3
- [Latent-Mark: An Audio Watermark Robust to Neural Codec Compression](chen26u_interspeech.md) — same problem · relatedness 2.2/3
- [Countering Neural Audio Codec Distortions in Watermarking with Adaptive Restoration](park26d_interspeech.md) — same problem · relatedness 2.1/3
- [FakeSound2: A Benchmark for Explainable, Traceable, and Generalizable Deepfake Sound Detection](xie26_interspeech.md) — shared data / evaluation · relatedness 2.1/3
- [DuraMark: Duration-Embedded Watermarking in LLM-based TTS](mou26_interspeech.md) — same problem · relatedness 2.1/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
