TL;DR — This paper introduces EO-LoRA and Flow-DGPO to achieve continuous, time-varying Valence-Arousal-Dominance (VAD) emotion control in zero-shot flow-matching text-to-speech without requiring massive emotional training corpora. The proposed approach outperforms prior baseline models on emotion similarity and prosody metrics while substantially reducing word error rates.
Key contributions
- Proposed Emotion Orthogonal LoRA (EO-LoRA), injecting three VAD-aligned low-rank branches into backbone linear layers with an orthogonality regularizer to prevent dimension interference.
- Introduced Flow-DGPO, a preference-based reinforcement learning alignment stage optimized via candidate group comparison to enhance emotional controllability while preserving speaker similarity and intelligibility.
- Demonstrated strong zero-shot time-varying emotion transitions and cross-lingual emotion transfer using only limited emotional training data (~41.2 hours total).
Problem
High-quality emotional text-to-speech models typically rely on coarse discrete emotion categories or massive proprietary emotion corpora, failing to support fine-grained intensity changes and intra-sentence time-varying transitions. Reference-based zero-shot prompting strategies entangle speaker identity, linguistic content, and emotional prosody into a single signal, making independent manipulation impossible. Existing methods like EmoCtrl-TTS require extensive data resources (over 60k hours), while standard LoRA updates conflate distinct emotional axes. Developing a data-efficient mechanism for fine-grained, independent, continuous, and time-varying emotional control remains an open challenge.
Method
The framework builds upon a pretrained F5-TTS Diffusion Transformer (DiT) flow-matching backbone. To achieve continuous emotion control, EO-LoRA injects parallel low-rank adaptation branches into selected linear layers. Specifically, for each emotion dimension , a low-rank update is defined as (where rank ), and scaled at every time step by the corresponding scalar component from the frame-level VAD sequence . Ablations indicate that injecting these branches into Attention Value and Output projections alongside Feed-Forward network linear layers yields optimal control stability. To prevent the three low-rank branches from learning correlated updates, an orthogonality regularizer computes cosine similarity via Frobenius inner products in an efficient trace form over matrices, using a regularization weight .
In the second stage, Flow-DGPO aligns the model using a direct preference optimization formulation adapted for flow-matching. Given a conditioning context, the model samples a group of candidates, computes composite rewards balancing emotion similarity (), speaker similarity (), and Word Error Rate (), and assigns GRPO-style normalized advantages. The policy is optimized using a preference margin objective against a frozen reference model with and 32 sampling steps, steering the conditional flow matching loss toward candidates that better express the target emotion trajectory without degrading intelligibility or speaker timbre.
Experimental setup
The models are trained on EmoVoice-DB (40 hours, 20k English samples) and the English subset of ESD (1.2 hours per speaker across 5 speakers/emotions, ~12 hours total). Evaluations are performed on two benchmarks: EMO-Change (RAVDESS-derived sentence pairs featuring explicit intra-utterance emotion shifts) and JVNV S2ST (Japanese-to-English cross-lingual emotion transfer). Baselines include the fine-tuned F5-TTS backbone and reported scores for EmoCtrl-TTS. Automatic metrics consist of Whisper-large-v3 WER, WavLM-large speaker similarity (SIM-o), AutoPCP multilingual v2, emotion2vec Emo SIM, and predictor Aro-Val SIM. Subjective evaluations utilize 5-point Likert Scale MOS (SMOS, NMOS, EMOS) evaluated by 20 native English listeners on 20 utterances.
Results
On the EMO-change test set, EO-LoRA combined with Flow-DGPO achieved an Emo SIM of 0.778, Aro-Val SIM of 0.914, AutoPCP of 3.60, and a low WER of 0.2%, outperforming the fine-tuned F5-TTS baseline (Emo SIM 0.693, Aro-Val SIM 0.837, AutoPCP 3.17, WER 7.1%). On the JVNV S2ST cross-lingual benchmark, EO-LoRA + Flow-DGPO reached an Emo SIM of 0.701 and Aro-Val SIM of 0.692, compared to 0.632 and 0.601 for the fine-tuned F5-TTS baseline. Ablations confirmed that removing the orthogonality regularizer () degraded Emo SIM from 0.753 to 0.708 on EMO-change, and injecting into Q/K/V/O projections caused optimization instabilities compared to the optimal Value, Output, and Feed-Forward (v+o+ff) configuration.
| System | SIM-o ↑ | WER (%) ↓ | AutoPCP ↑ | Emo SIM ↑ | Aro-Val SIM ↑ |
|---|---|---|---|---|---|
| F5TTS (fine-tuned) | 0.694 | 7.1 | 3.17 | 0.693 | 0.837 |
| EO-LoRA w/o Lorth | 0.721 | 0.6 | 3.31 | 0.708 | 0.843 |
| EO-LoRA | 0.741 | 0.6 | 3.51 | 0.753 | 0.905 |
| EO-LoRA + Flow-DGPO | 0.751 | 0.2 | 3.60 | 0.778 | 0.914 |
Limitations
The evaluation is restricted to English target speech and cross-lingual transfer from Japanese audio prompts, leaving multi-language scaling unexplored. The approach relies on an external pretrained wav2vec2 VAD predictor to extract frame-level continuous labels, meaning generation quality depends on the accuracy of the upstream emotion trajectory extractor. Compute requirements for group-sampling reinforcement learning (Flow-DGPO) are substantially higher than standard supervised fine-tuning.
Why read this
Researchers and engineers working on expressive text-to-speech and flow-matching alignment should read this paper to learn how to inject multi-dimensional, continuous control signals into pretrained backbones via orthogonal low-rank subspaces and preference optimization.
Code
Applications
Dynamic emotional text-to-speech for conversational AI assistants, video game character voice generation, audiobook narration with continuous affective shifts, and cross-lingual dubbing.
Institutions
Shanghai Jiao Tong University, Ping An Technology, Shanghai Jiao Tong University Chongqing Artificial Intelligence Research Institute
Related
- ETC-TTS: Emotion Trajectory Learning for Controllable Emotional Text-to-Speech — same problem · relatedness 2.9/3
- EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis — same problem · relatedness 2.8/3
- Emo-BPO: Emotion Bidirectional Preference Optimization for Diffusion-based Emotional TTS — same problem · relatedness 2.7/3
- Word-level Emotional Intensity Control in TTS via Emotion Residual Vectors — same problem · relatedness 2.6/3
- Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models — same problem · relatedness 2.5/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1798