All papers
Speech synthesisFull-paper digest

Continuous Time-Varying Emotion Control Zero-Shot Text-To-Speech With Emotion Orthogonal LoRA

Chenchen Wan, Minchuan Chen, Yan Shi, Peng Qi, Shaojun Wang, Jing Xiao

Code & resourceswancc-p.github.io/EoLoRA

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.4 KB · Ready to paste

Preview copied content

TL;DR — This paper introduces EO-LoRA and Flow-DGPO to achieve continuous, time-varying Valence-Arousal-Dominance (VAD) emotion control in zero-shot flow-matching text-to-speech without requiring massive emotional training corpora. The proposed approach outperforms prior baseline models on emotion similarity and prosody metrics while substantially reducing word error rates.

Key contributions

  • Proposed Emotion Orthogonal LoRA (EO-LoRA), injecting three VAD-aligned low-rank branches into backbone linear layers with an orthogonality regularizer to prevent dimension interference.
  • Introduced Flow-DGPO, a preference-based reinforcement learning alignment stage optimized via candidate group comparison to enhance emotional controllability while preserving speaker similarity and intelligibility.
  • Demonstrated strong zero-shot time-varying emotion transitions and cross-lingual emotion transfer using only limited emotional training data (~41.2 hours total).

Problem

High-quality emotional text-to-speech models typically rely on coarse discrete emotion categories or massive proprietary emotion corpora, failing to support fine-grained intensity changes and intra-sentence time-varying transitions. Reference-based zero-shot prompting strategies entangle speaker identity, linguistic content, and emotional prosody into a single signal, making independent manipulation impossible. Existing methods like EmoCtrl-TTS require extensive data resources (over 60k hours), while standard LoRA updates conflate distinct emotional axes. Developing a data-efficient mechanism for fine-grained, independent, continuous, and time-varying emotional control remains an open challenge.

Method

The framework builds upon a pretrained F5-TTS Diffusion Transformer (DiT) flow-matching backbone. To achieve continuous emotion control, EO-LoRA injects parallel low-rank adaptation branches into selected linear layers. Specifically, for each emotion dimension k∈{V,A,D}k \in \{V, A, D\}, a low-rank update is defined as ΔWk=AkBk\Delta W_k = A_k B_k (where rank r=16,α=32r=16, \alpha=32), and scaled at every time step tt by the corresponding scalar component from the frame-level VAD sequence (Vt,At,Dt)(V_t, A_t, D_t). Ablations indicate that injecting these branches into Attention Value and Output projections alongside Feed-Forward network linear layers yields optimal control stability. To prevent the three low-rank branches from learning correlated updates, an orthogonality regularizer computes cosine similarity via Frobenius inner products in an efficient trace form over r×rr \times r matrices, using a regularization weight λorth=0.05\lambda_{orth} = 0.05.

In the second stage, Flow-DGPO aligns the model using a direct preference optimization formulation adapted for flow-matching. Given a conditioning context, the model samples a group of G=8G=8 candidates, computes composite rewards balancing emotion similarity (λemo=3\lambda_{emo}=3), speaker similarity (λspk=1\lambda_{spk}=1), and Word Error Rate (λwer=1\lambda_{wer}=1), and assigns GRPO-style normalized advantages. The policy is optimized using a preference margin objective against a frozen reference model with β=20\beta=20 and 32 sampling steps, steering the conditional flow matching loss toward candidates that better express the target emotion trajectory without degrading intelligibility or speaker timbre.

Experimental setup

The models are trained on EmoVoice-DB (40 hours, 20k English samples) and the English subset of ESD (1.2 hours per speaker across 5 speakers/emotions, ~12 hours total). Evaluations are performed on two benchmarks: EMO-Change (RAVDESS-derived sentence pairs featuring explicit intra-utterance emotion shifts) and JVNV S2ST (Japanese-to-English cross-lingual emotion transfer). Baselines include the fine-tuned F5-TTS backbone and reported scores for EmoCtrl-TTS. Automatic metrics consist of Whisper-large-v3 WER, WavLM-large speaker similarity (SIM-o), AutoPCP multilingual v2, emotion2vec Emo SIM, and predictor Aro-Val SIM. Subjective evaluations utilize 5-point Likert Scale MOS (SMOS, NMOS, EMOS) evaluated by 20 native English listeners on 20 utterances.

Results

On the EMO-change test set, EO-LoRA combined with Flow-DGPO achieved an Emo SIM of 0.778, Aro-Val SIM of 0.914, AutoPCP of 3.60, and a low WER of 0.2%, outperforming the fine-tuned F5-TTS baseline (Emo SIM 0.693, Aro-Val SIM 0.837, AutoPCP 3.17, WER 7.1%). On the JVNV S2ST cross-lingual benchmark, EO-LoRA + Flow-DGPO reached an Emo SIM of 0.701 and Aro-Val SIM of 0.692, compared to 0.632 and 0.601 for the fine-tuned F5-TTS baseline. Ablations confirmed that removing the orthogonality regularizer (LorthL_{orth}) degraded Emo SIM from 0.753 to 0.708 on EMO-change, and injecting into Q/K/V/O projections caused optimization instabilities compared to the optimal Value, Output, and Feed-Forward (v+o+ff) configuration.

SystemSIM-o ↑WER (%) ↓AutoPCP ↑Emo SIM ↑Aro-Val SIM ↑
F5TTS (fine-tuned)0.6947.13.170.6930.837
EO-LoRA w/o Lorth0.7210.63.310.7080.843
EO-LoRA0.7410.63.510.7530.905
EO-LoRA + Flow-DGPO0.7510.23.600.7780.914

Limitations

The evaluation is restricted to English target speech and cross-lingual transfer from Japanese audio prompts, leaving multi-language scaling unexplored. The approach relies on an external pretrained wav2vec2 VAD predictor to extract frame-level continuous labels, meaning generation quality depends on the accuracy of the upstream emotion trajectory extractor. Compute requirements for group-sampling reinforcement learning (Flow-DGPO) are substantially higher than standard supervised fine-tuning.

Why read this

Researchers and engineers working on expressive text-to-speech and flow-matching alignment should read this paper to learn how to inject multi-dimensional, continuous control signals into pretrained backbones via orthogonal low-rank subspaces and preference optimization.

Code

Applications

Dynamic emotional text-to-speech for conversational AI assistants, video game character voice generation, audiobook narration with continuous affective shifts, and cross-lingual dubbing.

Institutions

Shanghai Jiao Tong University, Ping An Technology, Shanghai Jiao Tong University Chongqing Artificial Intelligence Research Institute

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1798