---
id: shang26_interspeech
title: "Seed-Enh: Generative Speech Enhancement in Decoupled Semantic and Timbre
  Spaces"
authors:
  - Zengqiang Shang
  - Biao Liu
  - Yu Zhao
  - Pengyuan Zhang
year: 2026
doi: 10.21437/Interspeech.2026-200
isca_url: https://www.isca-archive.org/interspeech_2026/shang26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/shang26_interspeech.pdf
session: Generative and Self-Supervised Speech Enhancement
topics:
  - speech-enhancement
  - voice-conversion
  - self-supervised
category: enhancement-separation
labels:
  - self-supervised
  - generative-model
  - robustness-noise
institutions:
  - Institute of Acoustics
  - University of Chinese Academy of Sciences
funding:
  - National Natural Science Foundation of China
  - CPSF Postdoctoral Fellowship
code:
  url: https://github.com/shangqwe123/Seed-Enh
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: shang26_interspeech
  category: enhancement-separation
  labels:
    - self-supervised
    - generative-model
    - robustness-noise
  institutions:
    - Institute of Acoustics
    - University of Chinese Academy of Sciences
  code: https://github.com/shangqwe123/Seed-Enh
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-200
  pdf: https://www.isca-archive.org/interspeech_2026/shang26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/shang26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/shang26_interspeech/markdown.md
---

# Seed-Enh: Generative Speech Enhancement in Decoupled Semantic and Timbre Spaces

*Zengqiang Shang, Biao Liu, Yu Zhao, Pengyuan Zhang*

[PDF](https://www.isca-archive.org/interspeech_2026/shang26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/shang26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-200)

**Category:** `enhancement-separation` · **Labels:** `self-supervised`, `generative-model`, `robustness-noise`

**TL;DR** — Seed-Enh is a generative speech enhancement framework that performs denoising in decoupled semantic and timbre spaces rather than traditional acoustic space, achieving a top DNS blind test OVRL of 3.095.

## Key contributions

- Extends decoupled space processing to speech enhancement via a three-stage pipeline: semantic extraction, timbre extraction, and flow matching fusion.
- Employs a frozen Whisper-Large-v2 encoder on an audio-splitting strategy to obtain noise-robust semantic representations without task-specific fine-tuning.
- Combines global CAM++ speaker embeddings with local context learning on Mel spectrograms for robust timbre reconstruction.
- Supports dual-mode execution, providing state-of-the-art speech enhancement and superior zero-shot voice conversion from noisy inputs.

## Problem

Most current speech enhancement systems operate directly in entangled acoustic spaces (such as STFT, Mel spectrograms, or time-domain waveforms), forcing models to learn shortcut solutions. This acoustic entanglement causes semantic models to either over-suppress speech or under-suppress noise, and limits timbre understanding leading to high-frequency attenuation, background holes, and speaker identity loss. Prior discriminative and generative baselines (like FullSubNet, TFGridNet, SGMSE, StoRM, and Schrödringer Bridge) either struggle with high-frequency restoration or create severe spectral artifacts. Decoupled generative frameworks used in TTS and voice conversion (like SeedTTS, SeedVC, and CosyVoice) have rarely been adapted successfully for noise-robust speech restoration.

## Method

The framework processes noisy speech through three decoupled stages. First, the input audio is split; the second half undergoes semantic extraction using a frozen Whisper-Large-v2 encoder (yielding 1024-dim representations downsampled by 4x), leveraging its pre-trained noise robustness to isolate linguistic content. Second, the first half of the audio is used for timbre extraction, combining a 192-dimensional global speaker embedding from a pre-trained CAM++ model with local context learning via concatenated Mel spectrograms and semantic features to capture fine-grained timbre details.

Third, information from both spaces is fused via a Diffusion Transformer (DiT) flow matching model. The DiT architecture concatenates upsampled semantic features and state representations with diffusion time steps and global timbre embeddings, utilizing adaptive layer normalization (AdaLN) and rotary position encoding (RoPE). Training follows a two-stage strategy: first pre-training on clean Emilia data to learn internal voice conversion mapping, and then fine-tuning on noisy-clean pairs (SNRs ranging from -5 to 15 dB).

During inference, clean Mel spectrograms are generated by solving the Ordinary Differential Equation (ODE) using an Euler solver with 25 steps. These generated spectrograms are subsequently converted into time-domain waveforms using a BigVGAN vocoder.

## Experimental setup

Models are trained on 101k hours of speech data from the Emilia dataset across multiple languages and 181 hours of noise from the DNS Challenge library (SNRs from -5 to 15 dB). Evaluation uses the DNS Challenge blind test set and simulated LibriTTS+wham dataset, measured via DNSMOS (OVRL, SIG, BAK), Resemblyzer SIM, and HuBERT-Large CER. All audio is resampled to 16,000 Hz, and inference employs a 25-step Euler ODE solver.

## Results

Seed-Enh achieves the highest overall quality scores on the DNS blind test set with an OVRL of 3.095 (vs. Schrödringer Bridge at 3.076, StoRM at 2.981, and FullSubNet at 2.772) and on LibriTTS+wham with an OVRL of 3.193. For voice conversion from noisy inputs, Seed-Enh significantly outperforms Seed-VC in overall quality (OVRL 3.225 vs. 2.461) and speaker similarity (SIM 0.823 vs. 0.726). While discriminative models achieve lower character error rates, Seed-Enh effectively balances noise suppression and harmonic restoration without suffering from the high-frequency attenuation or background holes observed in SGMSE and SB.

| Method | Type | DNS Blindtest OVRL ↑ | DNS Blindtest SIG ↑ | DNS Blindtest BAK ↑ | LibriTTS+wham OVRL ↑ | LibriTTS+wham SIM ↑ | LibriTTS+wham CER ↓ |
|---|---|---|---|---|---|---|---|
| FullSubNet | Regression | 2.772 | 3.134 | 3.755 | 2.587 | 0.835 | 0.128 |
| TFGridNet | Regression | 2.872 | 3.153 | 3.995 | 2.652 | 0.848 | 0.133 |
| SGMSE | Diffusion | 2.987 | 3.318 | 3.899 | 3.117 | 0.867 | 0.307 |
| SB | Schrödinger Bridge | 3.076 | 3.333 | 4.108 | 3.147 | 0.881 | 0.159 |
| Seed-Enh | Flow Matching | 3.095 | 3.361 | 4.071 | 3.193 | 0.890 | 0.167 |

## Limitations

The model relies on a frozen Whisper encoder and fixed hyperparameter settings without task-specific fine-tuning, which caps its maximum speech intelligibility gains relative to pure discriminative baselines. Inference requires a multi-step ODE solver (25 steps), rendering it computationally heavier than regression models. Furthermore, evaluation is primarily constrained to English and multi-language datasets at 16 kHz, leaving extreme low-resource dialects and high-sample-rate telephony untested.

## Why read this

Read this if you want to understand how to decouple semantic and timbre representations for generative audio restoration, or if you are building speech enhancement systems that require strong perceptual quality and zero-shot voice conversion capabilities.

## Code

- https://github.com/shangqwe123/Seed-Enh

## Applications

Robust real-time remote communication, speech pre-processing for automatic speech recognition, and zero-shot voice conversion under adverse acoustic environments.

## Institutions / 機構

Institute of Acoustics, University of Chinese Academy of Sciences

**Funding / 經費:** National Natural Science Foundation of China, CPSF Postdoctoral Fellowship

## Related

- [UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement](yan26_interspeech.md) — same problem · relatedness 3.0/3
- [HFMSE: Harmonic-Guided Speech Enhancement with Flow Matching](li26l_interspeech.md) — same problem · relatedness 2.9/3
- [PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement](gao26e_interspeech.md) — same problem · relatedness 2.9/3
- [Post-Training Speech Enhancement Language Models with Perceptual Rewards](berdo26_interspeech.md) — same problem · relatedness 2.9/3
- [Towards Robust Generative Speech Enhancement Using Vector Quantisation-Based Neural Audio Codec](zhao26i_interspeech.md) — same problem · relatedness 2.9/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
