All papers
Speech synthesisFull-paper digest

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

Jeongsoo Choi, Ji-Hoon Kim, Shujie Hu, Joon Son Chung

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — ProsoCodec is a prosody-oriented neural speech codec that models prosody as a conditional residual using text and speaker prefix conditioning, achieving state-of-the-art zero-shot voice conversion with a Word Error Rate of 4.45% and a source timbre leakage score (SIMs) of 0.167.

Key contributions

  • Proposes ProsoCodec, a generative speech codec that models residual prosody by conditioning both encoder and decoder on explicit text and speaker priors, bypassing complex adversarial disentanglement.
  • Introduces a dual-utterance training strategy using paired same-speaker utterances to decouple utterance-level prosody from global timbre and prevent prompt-style leakage.
  • Applies a low-frequency mel-band input restriction to bias the discrete codec bottleneck toward prosodic variation instead of fine-grained spectral details.
  • Demonstrates superior zero-shot voice conversion performance across objective and subjective metrics compared to established baselines like Vevo, Seed-VC, and FACodec.

Problem

Traditional neural speech codecs learn holistic representations that entangle linguistic content, speaker identity, and prosody, making them effective for zero-shot voice cloning but poor for voice conversion where source prosody must be strictly preserved. Prior approaches attempt to decompose speech or use restrictive bottlenecks and adversarial objectives, but prosody inherently depends on content and speaker context, meaning fully speaker-independent prosody representations destroy expressive nuance. Enforcing holistic imitation via prompt-based decoders also causes prompt-style leakage, overriding the target source prosody. This paper tackles these limitations to enable fine-grained prosodic control without sacrificing naturalness or target timbre adaptation.

Method

ProsoCodec combines an 8-layer Transformer encoder, a binary spherical quantizer (BSQ), and a 16-layer Diffusion Transformer (DiT) decoder, initialized from TaDiCodec with 1024 hidden dimensions, 4096 intermediate size, and 16 attention heads. Explicit priors are injected as prefix tokens by processing ASR transcripts through an MLP-based text encoder and extracting speaker embeddings via a pretrained speaker verification model (CAM++), concatenating them with input features as [e_spk; e_txt; e_mel].

To capture residual prosody, the continuous encoder representation undergoes linear interpolation and binary spherical quantization (BSQ) into an implicit binary codebook of size 4,096 at 12.5 Hz frame rate (150 bps bitrate). The input mel-spectrogram is restricted to its low-frequency band for the encoder to bias tokens toward prosody, while the decoder utilizes full-band mels with random span masking and flow-matching loss to handle conditional generation.

During training, a dual-utterance strategy alternates random span masking with paired same-speaker utterances (one as source, one as prompt) to break prompt-style copying. At inference, source speech tokens and transcript are fed alongside a reference prompt to generate the converted waveform via 32 Euler ODE steps and a Vocos vocoder.

Experimental setup

Trained on the LibriTTS dataset (585 hours of 24 kHz read speech, 2,456 speakers). Evaluated on LibriTTS test-clean, test-other, and VCTK splits using 1,000 randomly sampled 2-8 second utterances. Baselines include DDDM-VC, UniAudio, HierSpeech++, FACodec, Seed-VC, and Vevo. Metrics include WER (evaluated via Whisper-large-v3), SIMr (reference speaker similarity via WavLM-Large), SIMs (source timbre leakage), P-MOS, S-MOS, UTMOS, N-MOS, and log-scaled F0 RMSE. Models are trained for 150k updates with AdamW (learning rate 2e-5, global batch size of 160 seconds).

Results

ProsoCodec achieves a headline WER of 4.451%, outperforming Seed-VC (5.078%) and Vevo (4.826%), while scoring the best target speaker similarity (SIMr of 0.565) and lowest source timbre leakage (SIMs of 0.167). Ablation studies show that removing the low-frequency mel input degrades acoustic focus, while removing dual-utterance training raises F0 RMSE from 0.376 to 0.421, confirming worse prosody preservation. Removing text conditioning causes a catastrophic failure with WER exploding to 86.601%.

System / ConditionWER ↓SIMr ↑SIMs ↓RMSE (F0) ↓P-MOS ↑
Source Speech3.6680.0901.0000.000-
FACodec5.4540.3540.3470.4553.364
Seed-VC5.0780.5310.2390.4733.463
Vevo4.8260.4780.2430.4643.506
ProsoCodec (Ours)4.4510.5650.1670.4283.852

Limitations

The approach relies heavily on the accuracy of external pretrained ASR and speaker verification models to supply robust text and speaker priors. Evaluation is restricted to clean read speech corpora (LibriTTS and VCTK), leaving robustness to noisy or conversational speech unverified. The framework requires paired same-speaker training data or careful dual-utterance construction strategies.

Why read this

Researchers and engineers building zero-shot voice conversion or expressive speech generation systems should read this to learn how explicit conditioning and low-frequency spectral restriction can force a standard neural codec bottleneck to capture residual prosody without adversarial losses.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Zero-shot voice conversion, expressive text-to-speech, and spoken dialogue systems requiring precise prosody transfer across diverse speakers.

Institutions

KAIST, Chung-Ang University, Chinese University of Hong Kong

Funding / 經費: Institute of Information & Communications Technology Planning & Evaluation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2146