All papers
Resources & evaluationFull-paper digest

VoxEffects: A Speech-Oriented Audio Effects Dataset and Benchmark

Zhe Zhang, Yigitcan Özer, Junichi Yamagishi

Code & resourcesgithub.com/nii-yamagishilab/VoxEffects

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.4 KB · Ready to paste

Preview copied content

TL;DR — VoxEffects introduces a speech-oriented audio effects dataset and multi-task benchmark (AudioMAE-Fx) to infer applied post-production chains and parameters, achieving up to 95.58% macro presence accuracy under robust degradation training.

Key contributions

  • VoxEffects dataset featuring clean speech processed through a 6-stage canonical speech post-production chain with 2,520 curated preset combinations.
  • A reproducible audio renderer supporting offline synthesis and on-the-fly generation with multi-granularity supervision (presence, preset, count, and intensity).
  • A standardized robustness protocol simulating capture-side and platform-side degradations (noise, resampling, codecs) across in-domain and out-of-domain evaluation splits.
  • AudioMAE-Fx, an AudioMAE-based multi-task baseline demonstrating the necessity of curriculum-style robustness fine-tuning for cross-corpus generalization.

Problem

Real-world speech audio is routinely processed by post-production effects that improve intelligibility but shift signal statistics, confounding downstream systems, audio forensics, and content understanding. Existing audio effect identification (AEI) research largely targets music production regimes (e.g., guitar or singing vocals) rather than speech pipelines, or focuses on binary anti-spoofing tasks instead of attributing benign processing chains. Furthermore, prior work fails to evaluate robustness against real-world capture and platform artifacts like resampling ladders and lossy compression, leaving a major gap in speech production understanding.

Method

VoxEffects models a fixed canonical speech post-production chain consisting of six ordered effects: Denoising (DN), Dynamic Range Compression (DRC), Equalization (EQ), De-essing (DS), Reverberation (RVB), and Limiting (LIM), implemented via the Pedalboard library. Each effect e∈Ve \in \mathcal{V} is paired with a discrete preset bank PeP_e containing a bypass state and KeK_e quality-oriented presets (totaling 2,520 preset tuples p\mathbf{p}). Degradations D(⋅)D(\cdot) are independently parameterized via capture-side (DpreD_{\text{pre}}) and platform-side (DpostD_{\text{post}}) modules using additive noise, resampling, quantization, and lossy codecs under five settings: None, Pre-only, Post-only, Either, and Both.

The baseline model, AudioMAE-Fx, takes 16 kHz log-mel filterbank features as input and feeds them into a pretrained AudioMAE backbone. It employs lightweight prediction heads trained jointly via a multi-task objective: binary cross-entropy for KK-way effect presence (LpresL_{\text{pres}}, weighted by λpres=5\lambda_{\text{pres}}=5), cross-entropy for CC-way preset classification (C=2520C=2520), classification for active effect counts (L#actL_{\text{\#act}}), and L1 losses for scalar (LsL_s) and vector intensity regression (LvL_v).

Training proceeds in two stages: Stage 1 fine-tunes on clean rendered data using AdamW with a base learning rate of 10−310^{-3}, weight decay 0.050.05, and a layer-wise learning rate decay factor of 0.750.75 until plateau. Stage 2 executes robustness fine-tuning for an additional 50,000 steps by curriculum-style training on data augmented with 'Both' capture and platform degradations.

Experimental setup

The dataset is built from clean source corpora (DAPS, EARS, TSP) split 8:1:1 for train/val/test to evaluate in-domain (ID) performance, using VCTK for out-of-domain (OOD) generalization. Models are evaluated across five degradation configurations using fixed subsets of 60 ID and 60 OOD utterances rendered across all 2,520 presets. Metrics include macro-averaged accuracy (Accmacro\text{Acc}_{\text{macro}}), exact match ratio (EMR), Top-1/Top-5 preset accuracy, active count accuracy, and mean absolute error (MAEmean\text{MAE}_{\text{mean}}, MAEoverall\text{MAE}_{\text{overall}}) for intensity regression.

Results

Robustness fine-tuning substantially outperforms baseline training without augmentation, boosting OOD effect presence detection accuracy from 82.81% to 86.15% (and macro accuracy from 91.59% to 95.58% ID under 'Both' conditions). For fine-grained preset classification (2,520 classes), Top-1 accuracy improves from 5.76% to 12.19% on OOD and 21.52% to 36.78% ID when using degradation augmentation, though absolute numbers remain challenging due to perceptual overlap. In intensity regression, degradation fine-tuning reduces mean vector intensity MAE from 0.22 to 0.19 on OOD test sets.

Test AugmentationTrain AugmentationEffect Presence Acc_macro (ID/OOD)Exact Match Ratio (ID/OOD)Preset Top-1 Acc. (ID/OOD)#Active Acc. (ID/OOD)Intensity MAE_mean (ID/OOD)
NoneNone91.59 / 82.8158.96 / 30.8621.52 / 5.7661.11 / 45.810.14 / 0.22
NoneBoth95.58 / 86.1576.48 / 39.2236.78 / 12.1977.24 / 47.360.10 / 0.19
BothNone75.42 / 71.1321.68 / 13.854.54 / 1.7640.72 / 39.850.27 / 0.31
BothBoth88.48 / 80.8749.77 / 27.5812.57 / 5.4856.57 / 39.780.17 / 0.23

Limitations

The dataset assumes a rigid, fixed post-production chain ordering and a finite preset bank, ignoring alternative orderings, repeated processing stages, or continuous parameter variations. The rendering pipeline relies exclusively on the Pedalboard library, risking implementation mismatch with commercial toolchains. Certain effects like conservative denoising and limiters exhibit weak alignments between preset labels and observable acoustic artifacts under domain shifts. Finally, evaluation is currently restricted to fixed-length segments and a single AudioMAE backbone architecture.

Why read this

Researchers and engineers tackling speech forensics, content understanding, or audio engineering assistance should read this paper to adopt the first standardized speech-oriented audio effects benchmark and understand how data degradation impacts multi-granularity effect inference.

Code

Applications

Audio forensics, production-aware speech content understanding, automated audio engineering assistance, and educational ear-training tools for sound engineers.

Institutions

National Institute of Informatics

Funding / 經費: New Energy and Industrial Technology Development Organization

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1621