TL;DR — VoxEffects introduces a speech-oriented audio effects dataset and multi-task benchmark (AudioMAE-Fx) to infer applied post-production chains and parameters, achieving up to 95.58% macro presence accuracy under robust degradation training.
Key contributions
- VoxEffects dataset featuring clean speech processed through a 6-stage canonical speech post-production chain with 2,520 curated preset combinations.
- A reproducible audio renderer supporting offline synthesis and on-the-fly generation with multi-granularity supervision (presence, preset, count, and intensity).
- A standardized robustness protocol simulating capture-side and platform-side degradations (noise, resampling, codecs) across in-domain and out-of-domain evaluation splits.
- AudioMAE-Fx, an AudioMAE-based multi-task baseline demonstrating the necessity of curriculum-style robustness fine-tuning for cross-corpus generalization.
Problem
Real-world speech audio is routinely processed by post-production effects that improve intelligibility but shift signal statistics, confounding downstream systems, audio forensics, and content understanding. Existing audio effect identification (AEI) research largely targets music production regimes (e.g., guitar or singing vocals) rather than speech pipelines, or focuses on binary anti-spoofing tasks instead of attributing benign processing chains. Furthermore, prior work fails to evaluate robustness against real-world capture and platform artifacts like resampling ladders and lossy compression, leaving a major gap in speech production understanding.
Method
VoxEffects models a fixed canonical speech post-production chain consisting of six ordered effects: Denoising (DN), Dynamic Range Compression (DRC), Equalization (EQ), De-essing (DS), Reverberation (RVB), and Limiting (LIM), implemented via the Pedalboard library. Each effect is paired with a discrete preset bank containing a bypass state and quality-oriented presets (totaling 2,520 preset tuples ). Degradations are independently parameterized via capture-side () and platform-side () modules using additive noise, resampling, quantization, and lossy codecs under five settings: None, Pre-only, Post-only, Either, and Both.
The baseline model, AudioMAE-Fx, takes 16 kHz log-mel filterbank features as input and feeds them into a pretrained AudioMAE backbone. It employs lightweight prediction heads trained jointly via a multi-task objective: binary cross-entropy for -way effect presence (, weighted by ), cross-entropy for -way preset classification (), classification for active effect counts (), and L1 losses for scalar () and vector intensity regression ().
Training proceeds in two stages: Stage 1 fine-tunes on clean rendered data using AdamW with a base learning rate of , weight decay , and a layer-wise learning rate decay factor of until plateau. Stage 2 executes robustness fine-tuning for an additional 50,000 steps by curriculum-style training on data augmented with 'Both' capture and platform degradations.
Experimental setup
The dataset is built from clean source corpora (DAPS, EARS, TSP) split 8:1:1 for train/val/test to evaluate in-domain (ID) performance, using VCTK for out-of-domain (OOD) generalization. Models are evaluated across five degradation configurations using fixed subsets of 60 ID and 60 OOD utterances rendered across all 2,520 presets. Metrics include macro-averaged accuracy (), exact match ratio (EMR), Top-1/Top-5 preset accuracy, active count accuracy, and mean absolute error (, ) for intensity regression.
Results
Robustness fine-tuning substantially outperforms baseline training without augmentation, boosting OOD effect presence detection accuracy from 82.81% to 86.15% (and macro accuracy from 91.59% to 95.58% ID under 'Both' conditions). For fine-grained preset classification (2,520 classes), Top-1 accuracy improves from 5.76% to 12.19% on OOD and 21.52% to 36.78% ID when using degradation augmentation, though absolute numbers remain challenging due to perceptual overlap. In intensity regression, degradation fine-tuning reduces mean vector intensity MAE from 0.22 to 0.19 on OOD test sets.
| Test Augmentation | Train Augmentation | Effect Presence Acc_macro (ID/OOD) | Exact Match Ratio (ID/OOD) | Preset Top-1 Acc. (ID/OOD) | #Active Acc. (ID/OOD) | Intensity MAE_mean (ID/OOD) |
|---|---|---|---|---|---|---|
| None | None | 91.59 / 82.81 | 58.96 / 30.86 | 21.52 / 5.76 | 61.11 / 45.81 | 0.14 / 0.22 |
| None | Both | 95.58 / 86.15 | 76.48 / 39.22 | 36.78 / 12.19 | 77.24 / 47.36 | 0.10 / 0.19 |
| Both | None | 75.42 / 71.13 | 21.68 / 13.85 | 4.54 / 1.76 | 40.72 / 39.85 | 0.27 / 0.31 |
| Both | Both | 88.48 / 80.87 | 49.77 / 27.58 | 12.57 / 5.48 | 56.57 / 39.78 | 0.17 / 0.23 |
Limitations
The dataset assumes a rigid, fixed post-production chain ordering and a finite preset bank, ignoring alternative orderings, repeated processing stages, or continuous parameter variations. The rendering pipeline relies exclusively on the Pedalboard library, risking implementation mismatch with commercial toolchains. Certain effects like conservative denoising and limiters exhibit weak alignments between preset labels and observable acoustic artifacts under domain shifts. Finally, evaluation is currently restricted to fixed-length segments and a single AudioMAE backbone architecture.
Why read this
Researchers and engineers tackling speech forensics, content understanding, or audio engineering assistance should read this paper to adopt the first standardized speech-oriented audio effects benchmark and understand how data degradation impacts multi-granularity effect inference.
Code
Applications
Audio forensics, production-aware speech content understanding, automated audio engineering assistance, and educational ear-training tools for sound engineers.
Institutions
National Institute of Informatics
Funding / 經費: New Energy and Industrial Technology Development Organization
Related
- Audio-Visual Feature Reconstruction Pretraining for Noise-Robust Emotion Recognition — complementary · relatedness 1.8/3
- Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs — shared data / evaluation · relatedness 1.7/3
- NoiseLoRA-SV: Hierarchical Noise-Conditioned Adaptation with Embedding Distillation for Robust Speaker Verification — complementary · relatedness 1.7/3
- Robust Audio-Visual Emotion Recognition via Conditional Transformer U-Nets with Frequency-Injected Visual Stream — complementary · relatedness 1.7/3
- Coco-VC: Degradation-Robust Streaming Voice Conversion System on the Listener Side — complementary · relatedness 1.7/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1621