All papers
Speech synthesisFull-paper digest

Refining Emphasis Control in Flow-Matching TTS via Preference Alignment and Reinforcement Learning

Jiangnan Ye, Jiawei Jin, Pengfei Tan, Chao Yan, Xuerui Yang, Zhiyong Wu

Code & resourcesthuhcsi.github.io/interspeech2026-F5Emphasis

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.5 KB · Ready to paste

Preview copied content

TL;DR — This paper presents an emphasis-controllable non-autoregressive text-to-speech framework built by extending F5-TTS with a dedicated emphasis encoder and a three-stage training pipeline (SFT, DPO, and Flow-CPS reinforcement learning). The final model achieves a top WPT prominence score of 1.45 while maintaining a low Word Error Rate (1.63%) and high speaker similarity.

Key contributions

  • Proposed an emphasis-controllable F5-TTS architecture integrating an explicit Emphasis Encoder with 4 Transformer layers into the DiT backbone.
  • Introduced a three-stage optimization pipeline combining Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO) adapted for flow-matching.
  • Utilized the Wavelet Prosody Toolkit (WPT) to automatically compute objective prominence scores (pitch, energy, duration) as rewards for preference and RL alignment.
  • Demonstrated superior generalization on underrepresented categories like nouns, achieving 63.00% accuracy compared to 37.00% for F5-TTS-dpo and 22.00% for CosyVoice.

Problem

Fine-grained emphasis control remains a major bottleneck in neural text-to-speech due to data scarcity and the complexity of non-linear human prosody. Prior rule-based acoustic feature manipulation (e.g., SpeechCraft, WhiStress) yields robotic and unnatural speech, whereas automated language tools fail to capture subtle human perceptual nuances. Meanwhile, instruction-following large audio models and LLM-based TTS systems (like CosyVoice or Doubao) lack the precision and data efficiency needed for robust, high-intensity prosodic modulation under low-resource constraints.

Method

The framework extends the F5-TTS v1 Base model (DiT backbone with 1024 hidden dimension, depth 22, 16 attention heads, and ff multiplier 2) by adding an Emphasis Encoder consisting of 4 Transformer layers that share the hidden size and attention configuration of the base model. Input texts with emphasis markers (<strong> and </strong>) are processed through rotary positional embeddings, self-attention layers with residual connections, and feed-forward networks with layer normalization to generate emphasis representations. These are element-wise added to the original text token embeddings and supplied as conditioning signals to the DiT backbone.

The training pipeline follows three progressive stages: (1) Supervised Fine-Tuning on 2.5 hours of manually annotated Chinese speech data using the AdamW optimizer with a learning rate of 1×10⁻⁵ for 21,500 steps across 8 NVIDIA H800 (80GB) GPUs. (2) Direct Preference Optimization (DPO) on an 800-pair preference dataset generated by the SFT model and ranked using WPT prominence scores derived from Montreal Forced Aligned boundaries; optimized for 13,000 steps with an SFT loss weight of 0.5, DPO loss weight of 1.0, and temperature τ = 0.5. (3) Online reinforcement learning via Flow-CPS (a variant of FlowGRPO adapted for flow-matching), which uses WPT-based group-relative advantages to update latent trajectories while avoiding numerical instabilities via modified KL divergence penalties.

During inference, the model takes text prompts with explicit emphasis tags and generates flow-predicted target mel-spectrograms that exhibit targeted acoustic prominence without requiring runtime heuristic parameter tuning.

Experimental setup

The models were trained and evaluated on Chinese speech datasets, utilizing 2.5 hours of manually annotated data for SFT, 800 pairs for DPO, and 2,000 samples for GRPO. Evaluations were conducted on the SEED test-zh dataset for intelligibility and speaker similarity, alongside a dedicated test set of 200 DeepSeek-annotated samples for prominence, and a 100-sentence noun-emphasis test set. Baselines included the vanilla F5-TTS base model and CosyVoice. Metrics comprised Word Error Rate (WER), Speaker Identity Match (SIM) via WavLM, Wavelet Prosody Toolkit (WPT) Prominence, and subjective MOS evaluations (E-MOS and N-MOS) rated by 15 native Mandarin speakers.

Results

The progressive training stages steadily increase WPT prominence from 0.86 (vanilla F5-TTS) to 1.22 (F5-TTS-sft), 1.42 (F5-TTS-dpo), and finally 1.45 (F5-TTS-grpo), outperforming the CosyVoice baseline score of 1.32. Crucially, the proposed method preserves intelligibility and voice identity: WER remains stable between 1.62% and 1.63% across all F5 variants (compared to CosyVoice's higher WER of 3.63%), and SIM stays high at ~0.71–0.72. In subjective evaluations, F5-TTS-grpo outperforms CosyVoice in both Emphasis MOS (2.51 vs 2.16) and Naturalness MOS (3.47 vs 3.06). On challenging underrepresented noun tests, F5-TTS-grpo achieves an accuracy of 63.00% with a prominence of 1.28, vastly surpassing F5-TTS-dpo (37.00% ACC, 0.96 prominence) and CosyVoice (22.00% ACC, 0.83 prominence).

ModelProminence↑WER↓SIM↑
Cosyvoice1.323.63%0.72
F5-TTS0.861.56%0.74
F5-TTS-sft1.221.62%0.72
F5-TTS-dpo1.421.62%0.72
F5-TTS-grpo1.451.63%0.71

Limitations

The framework's scope is currently bounded by its evaluation primarily on Mandarin Chinese datasets and a relatively narrow domain of manually curated and DeepSeek-annotated training data (2.5 hours for SFT, 800 preference pairs, and 2,000 RL samples). The reliance on the Wavelet Prosody Toolkit and Montreal Forced Aligner for automated reward modeling means potential alignment errors or acoustic noise in difficult acoustic environments could degrade reward signal quality. Furthermore, compute requirements are high, needing clusters of 8 NVIDIA H800 GPUs for multi-stage training.

Why read this

Researchers and engineers working on controllable speech synthesis and preference alignment for non-autoregressive flow-matching models will find this a blueprint for adapting DPO and GRPO to continuous audio generation. It provides concrete recipes for bypassing the data-hunger of Audio LLMs through modular explicit encoders and stable reward-driven trajectory optimization.

Code

Applications

Expressive audiobook narration, conversational voice assistants requiring precise focus and intent modulation, and dynamic character dubbing in gaming and animation.

Institutions

Tsinghua University, StepFun

Funding / 經費: National Natural Science Foundation of China, National Social Science Foundation of China

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2284