TL;DR — Seed-Enh is a generative speech enhancement framework that performs denoising in decoupled semantic and timbre spaces rather than traditional acoustic space, achieving a top DNS blind test OVRL of 3.095.
Key contributions
- Extends decoupled space processing to speech enhancement via a three-stage pipeline: semantic extraction, timbre extraction, and flow matching fusion.
- Employs a frozen Whisper-Large-v2 encoder on an audio-splitting strategy to obtain noise-robust semantic representations without task-specific fine-tuning.
- Combines global CAM++ speaker embeddings with local context learning on Mel spectrograms for robust timbre reconstruction.
- Supports dual-mode execution, providing state-of-the-art speech enhancement and superior zero-shot voice conversion from noisy inputs.
Problem
Most current speech enhancement systems operate directly in entangled acoustic spaces (such as STFT, Mel spectrograms, or time-domain waveforms), forcing models to learn shortcut solutions. This acoustic entanglement causes semantic models to either over-suppress speech or under-suppress noise, and limits timbre understanding leading to high-frequency attenuation, background holes, and speaker identity loss. Prior discriminative and generative baselines (like FullSubNet, TFGridNet, SGMSE, StoRM, and Schrödringer Bridge) either struggle with high-frequency restoration or create severe spectral artifacts. Decoupled generative frameworks used in TTS and voice conversion (like SeedTTS, SeedVC, and CosyVoice) have rarely been adapted successfully for noise-robust speech restoration.
Method
The framework processes noisy speech through three decoupled stages. First, the input audio is split; the second half undergoes semantic extraction using a frozen Whisper-Large-v2 encoder (yielding 1024-dim representations downsampled by 4x), leveraging its pre-trained noise robustness to isolate linguistic content. Second, the first half of the audio is used for timbre extraction, combining a 192-dimensional global speaker embedding from a pre-trained CAM++ model with local context learning via concatenated Mel spectrograms and semantic features to capture fine-grained timbre details.
Third, information from both spaces is fused via a Diffusion Transformer (DiT) flow matching model. The DiT architecture concatenates upsampled semantic features and state representations with diffusion time steps and global timbre embeddings, utilizing adaptive layer normalization (AdaLN) and rotary position encoding (RoPE). Training follows a two-stage strategy: first pre-training on clean Emilia data to learn internal voice conversion mapping, and then fine-tuning on noisy-clean pairs (SNRs ranging from -5 to 15 dB).
During inference, clean Mel spectrograms are generated by solving the Ordinary Differential Equation (ODE) using an Euler solver with 25 steps. These generated spectrograms are subsequently converted into time-domain waveforms using a BigVGAN vocoder.
Experimental setup
Models are trained on 101k hours of speech data from the Emilia dataset across multiple languages and 181 hours of noise from the DNS Challenge library (SNRs from -5 to 15 dB). Evaluation uses the DNS Challenge blind test set and simulated LibriTTS+wham dataset, measured via DNSMOS (OVRL, SIG, BAK), Resemblyzer SIM, and HuBERT-Large CER. All audio is resampled to 16,000 Hz, and inference employs a 25-step Euler ODE solver.
Results
Seed-Enh achieves the highest overall quality scores on the DNS blind test set with an OVRL of 3.095 (vs. Schrödringer Bridge at 3.076, StoRM at 2.981, and FullSubNet at 2.772) and on LibriTTS+wham with an OVRL of 3.193. For voice conversion from noisy inputs, Seed-Enh significantly outperforms Seed-VC in overall quality (OVRL 3.225 vs. 2.461) and speaker similarity (SIM 0.823 vs. 0.726). While discriminative models achieve lower character error rates, Seed-Enh effectively balances noise suppression and harmonic restoration without suffering from the high-frequency attenuation or background holes observed in SGMSE and SB.
| Method | Type | DNS Blindtest OVRL ↑ | DNS Blindtest SIG ↑ | DNS Blindtest BAK ↑ | LibriTTS+wham OVRL ↑ | LibriTTS+wham SIM ↑ | LibriTTS+wham CER ↓ |
|---|---|---|---|---|---|---|---|
| FullSubNet | Regression | 2.772 | 3.134 | 3.755 | 2.587 | 0.835 | 0.128 |
| TFGridNet | Regression | 2.872 | 3.153 | 3.995 | 2.652 | 0.848 | 0.133 |
| SGMSE | Diffusion | 2.987 | 3.318 | 3.899 | 3.117 | 0.867 | 0.307 |
| SB | Schrödinger Bridge | 3.076 | 3.333 | 4.108 | 3.147 | 0.881 | 0.159 |
| Seed-Enh | Flow Matching | 3.095 | 3.361 | 4.071 | 3.193 | 0.890 | 0.167 |
Limitations
The model relies on a frozen Whisper encoder and fixed hyperparameter settings without task-specific fine-tuning, which caps its maximum speech intelligibility gains relative to pure discriminative baselines. Inference requires a multi-step ODE solver (25 steps), rendering it computationally heavier than regression models. Furthermore, evaluation is primarily constrained to English and multi-language datasets at 16 kHz, leaving extreme low-resource dialects and high-sample-rate telephony untested.
Why read this
Read this if you want to understand how to decouple semantic and timbre representations for generative audio restoration, or if you are building speech enhancement systems that require strong perceptual quality and zero-shot voice conversion capabilities.
Code
Applications
Robust real-time remote communication, speech pre-processing for automatic speech recognition, and zero-shot voice conversion under adverse acoustic environments.
Institutions
Institute of Acoustics, University of Chinese Academy of Sciences
Funding / 經費: National Natural Science Foundation of China, CPSF Postdoctoral Fellowship
Related
- UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement — same problem · relatedness 3.0/3
- HFMSE: Harmonic-Guided Speech Enhancement with Flow Matching — same problem · relatedness 2.9/3
- PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement — same problem · relatedness 2.9/3
- Post-Training Speech Enhancement Language Models with Perceptual Rewards — same problem · relatedness 2.9/3
- Towards Robust Generative Speech Enhancement Using Vector Quantisation-Based Neural Audio Codec — same problem · relatedness 2.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-200