All papers
Speech synthesisFull-paper digest

Word-level Emotional Intensity Control in TTS via Emotion Residual Vectors

Ji-Hyun Park, Nam-Seok Song, Joon-Hyuk Chang

Code & resourcesjjhh0210.github.io/EmoRes-tts-demo

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.0 KB · Ready to paste

Preview copied content

TL;DR — This paper introduces emotion residual vectors (ERVs) derived from self-supervised speech embeddings to enable word-level emotional intensity control in TTS while preserving naturalness, achieving an A/B preference win over baseline HED-TTS of 87.1%.

Key contributions

  • Proposed emotion residual vectors (ERVs) computed as neutral-to-emotional word-aligned deviations in self-supervised speech representation space as annotation-free local prosody cues.
  • Designed a residual injection module (RIM) with a low-dimensional bottleneck subspace to stabilize intensity modulation when added to a neutral-conditioned FastSpeech2 backbone.
  • Introduced a three-stage training pipeline integrating ERV extraction, RIM training with fixed neutral global conditioning, and a fine-tuned RoBERTa-based projected-ERV predictor.
  • Demonstrated continuous word-level emotional intensity scaling that avoids abrupt pitch shifts and maintains perceptual speech naturalness.

Problem

Traditional emotional text-to-speech models typically rely on utterance-level conditioning, which often causes prosodic misallocation where non-target words are incorrectly emphasized and intended focus words are attenuated. Prior word-level or fine-grained control schemes—such as EmoQ-TTS, HED-TTS, and EME-TTS—allow per-word adjustment of emotional strength but frequently introduce perceptual degradations, including abrupt pitch shifts, unnatural word durations, and overly salient target words. Maintaining perceptual naturalness while achieving precise local control remains a significant challenge due to emotion-emphasis interference and local prosodic distortions.

Method

The framework utilizes a three-stage pipeline built around a FastSpeech2 backbone (4 FFT blocks, hidden size 256, 2 attention heads, filter size 1024, kernel sizes 9 and 1) with an iSTFTNet vocoder (C8C8I architecture, hop length 192). First, the backbone is trained with speaker and utterance-level emotion conditioning. Second, the backbone is frozen, and a Residual Injection Module (RIM) is trained to map word-level ERVs into encoder hidden states via linear projections and LayerNorm, using a bottleneck dimension of a=32. ERVs are extracted by computing word-level mean-pooled WavLM-Base representations (aggregating layers 7 to 12 into 768-dimensional vectors) from parallel neutral-emotional word pairs aligned via the Montreal Forced Aligner on the ESD dataset. Third, a RoBERTa-Base predictor is trained for 20 epochs using AdamW (learning rates 5e-5 for RoBERTa and 3e-4 for the prediction head) with a word-wise MSE objective to estimate projected ERVs from text and emotion prompt embeddings.

During inference, word-level projected ERVs generated by the RoBERTa predictor are converted into additive hidden-state offsets through the frozen RIM, scaled by non-negative word-level control weights alpha, and injected into the neutral-conditioned FastSpeech2 encoder. This design isolates the neutral-to-emotional acoustic shift into a compact subspace, preventing the instability and quality degradation typical of high-dimensional direct residual manipulation.

Experimental setup

Experiments utilized the English subset of the Emotional Speech Database (ESD), featuring 10 speakers and 5 emotions across a train/validation/test split of 14,900/950/1,550 utterances, yielding 88,400 word-level ERV instances. The system was compared against utterance-level FastSpeech2 (FS2+emo), EME-TTS, and HED-TTS. Evaluation metrics include Naturalness Mean Opinion Score (NMOS), Emotional Mean Opinion Score (EMOS), UTMOS, word error rate (WER) via Whisper, and emotion classification accuracy (Emo.Acc.) using Emotion2vec-plus-large, alongside A/B preference listening tests.

Results

The proposed model achieves competitive speech quality and emotion classification accuracy compared to the unconstrained utterance-level FS2+emo baseline (NMOS 3.83 vs 3.81, Emo.Acc. 0.71 vs 0.70), while substantially outperforming specialized word-level baselines like HED-TTS (UTMOS 3.77 vs 2.90, Emo.Acc. 0.71 vs 0.54). In A/B listening tests for word-level emphasis naturalness, listeners preferred the proposed method over HED-TTS in 87.1% of trials (with only 5.4% losses) and over EME-TTS in 55.3% of trials. Ablation studies on the bottleneck dimension revealed that a bottleneck of a=32 provides an optimal trade-off, balancing word error rate (8.22) and emotion classification accuracy (0.62 for injection-only, rising to 0.71 with the full predictor).

ModelNMOS ↑EMOS ↑UTMOS ↑Emo.Acc. ↑
GT4.73 ± 0.154.04 ± 0.194.17 ± 0.030.72
iSTFTNet4.56 ± 0.174.21 ± 0.174.11 ± 0.030.74
FS2+emo3.81 ± 0.193.61 ± 0.213.70 ± 0.040.70
EME-TTS3.78 ± 0.193.50 ± 0.203.76 ± 0.030.71
HED-TTS2.75 ± 0.192.97 ± 0.212.90 ± 0.040.54
Proposed3.83 ± 0.193.56 ± 0.213.77 ± 0.040.71

Limitations

The approach relies strictly on parallel, word-aligned neutral-emotional pairs for extracting ERV targets, which restricts its direct application to unaligned or in-the-wild expressive speech data. Evaluation was limited to the English subset of a single controlled dataset (ESD), leaving cross-lingual generalizability and performance on spontaneous conversational speech unverified.

Why read this

Speech synthesis researchers working on fine-grained prosody and emotion control should read this to see how self-supervised feature residuals can be compressed into a stable bottleneck space for localized text-to-speech modulation without sacrificing acoustic naturalness.

Code

Applications

Interactive voice assistants, audiobook narration, and expressive character dubbing in video games requiring precise word-level emotional emphasis.

Institutions

Hanyang University

Funding / 經費: Institute of Information & Communications Technology Planning & Evaluation, Ministry of Science and ICT

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3079