All papers
Speech synthesisFull-paper digest

Learning to Rescale: On-the-Fly Sequence Length Adaptation in Non-Autoregressive Speech Synthesis

Jiawei Jin, Ren Wang, Zhiyu Cui, Shun Lei, Yixuan Zhou, Zhiyong Wu

Code & resourcesthuhcsi.github.io/ElasticDLM

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.7 KB · Ready to paste

Preview copied content

TL;DR — ElasticDLM is a variable-length non-autoregressive diffusion language model for text-to-speech that eliminates reliance on accurate pre-calculated character-level durations by introducing learnable functional tokens and dynamic length scaling, matching or exceeding baseline synthesis quality even under severe initial length mispredictions.

Key contributions

  • Proposes a variable-length non-autoregressive speech synthesis architecture integrating two functional tokens ([EXPAND] and [DELETE]) for unified, learnable length regulation.
  • Develops the Differentiated Length-Scaling Training Scheme (DLTS), combining block-wise merging/insertion with an auxiliary length-scaling loss to handle acoustic redundancy.
  • Introduces the Hierarchical Confidence-Guided Inference (HCGI) strategy for coarse-to-fine sequence length refinement alongside content generation with negligible latency overhead.
  • Achieves robust performance across extreme initialization variations (1s to 16s) and highly expressive emotional styles without requiring precise forced alignments.

Problem

Non-autoregressive (NAR) text-to-speech models typically rely on pre-extracted character-level durations or fixed global duration estimates from external aligners (like Montreal Forced Aligner or teacher-forcing models). Inaccurate duration predictions lead to severe audible artifacts such as unnatural speaking rates, fragmented prosody, and total synthesis failure under in-the-wild conditions. While variable-length diffusion language models exist in text generation (e.g., DreamOn), directly applying single-token insertion and deletion operations fails on speech due to high temporal resolution, acoustic redundancy, and strict inference budgets.

Method

ElasticDLM builds upon the 16-layer Transformer architecture (hidden size 1536, 16 attention heads) of MaskGCT's Text-to-Semantic model, taking phoneme tokens and timbre-prompt semantic tokens as condition input CC. The framework incorporates [EXPAND] and [DELETE] functional tokens into a cascaded Diffusion Language Model (DLM) to handle sequence length adjustments dynamically. During training, the Differentiated Length-Scaling Training Scheme (DLTS) applies standard content masking followed by a stochastic block-wise compression where non-overlapping blocks of nn masked tokens are merged with probability pp, and insertion counts are sampled uniformly from k∼Uniform(1,⌊α(t)⋅L⌋)k \sim \text{Uniform}(1, \lfloor \alpha(t) \cdot L \rfloor). The model is optimized using a composite objective combining standard negative log-likelihood (NLL) loss over masked positions and an auxiliary length loss weighted by λ\lambda that minimizes absolute error between ground-truth and predicted total length changes.

During inference, the Hierarchical Confidence-Guided Inference (HCGI) strategy processes sequences in a coarse-to-fine manner. At each step tt, the model extracts softmax confidence scores pexpp_{\text{exp}} and pdelp_{\text{del}} for the functional tokens; if they exceed threshold τ\tau (set to 0.8), operations are immediately triggered, executing [DELETE] removals or replacing [EXPAND] tokens with nn [MASK] tokens. Simultaneously, content logits utilize top-kk sampling (k=20k=20, temperature linearly annealed from 1.5 to 0 over 50 steps), and the lowest confidence tokens are re-masked according to a sine scheduling function.

Experimental setup

The model is trained on 100k hours of data from the English and Chinese subsets of Emilia-large using 24 NVIDIA A100 (80GB) GPUs. Training uses the AdamW optimizer with a 1e-4 learning rate and 32K warmup steps, setting expansion block size n=5n=5. Evaluations are conducted on the seed-test-zh and seed-testen benchmarks from Seed-TTS, comparing against MaskGCT variants (Normal, Fix with 4s/4.5s initialization, and GT oracle), E2-TTS, and F5-TTS, using Word Error Rate (WER), Len-L1 sequence length distance, WavLM cosine speaker similarity (Sim), and Mean Opinion Score (N-MOS).

Results

On the seed-test-zh benchmark, Ours-Normal achieves a WER of 2.387%, Len-L1 of 0.7091s, N-MOS of 4.20, and speaker similarity of 0.772, performing comparably to the oracle MaskGCT-GT while outperforming standard MaskGCT-Normal on naturalness and length accuracy. Under the challenging Fix setting (with fixed 4s/4.5s initialization), the baseline MaskGCT-Fix degrades severely (WER 4.182% zh / 13.341% en, Len-L1 1.6993s zh), whereas ElasticDLM (Ours-Fix) maintains strong robustness with a WER of 2.582% (zh) / 4.349% (en) and Len-L1 of 0.8752s (zh). Ablations confirm that removing the length-loss or HCGI strategy causes substantial performance drops, with w/o HCGI-Fix skyrocketing WER to 4.005% (zh) and 5.652% (en). In extreme initial length tests ranging from 1s to 16s, baseline models like E2-TTS, F5-TTS, and MaskGCT experience catastrophic WER degradation (up to 107.76% at 16s), whereas ElasticDLM constantly maintains low WER (~2.4% to 2.6%).

SystemWER (zh)Len-L1 (zh)N-MOS (zh)WER (en)Len-L1 (en)N-MOS (en)
MaskGCT-GT2.286%-4.23 ± 0.133.833%-4.27 ± 0.14
MaskGCT-Fix4.182%1.6993s3.72 ± 0.1613.341%1.0156s3.58 ± 0.19
MaskGCT-Normal2.330%0.8555s3.84 ± 0.183.936%0.5429s3.96 ± 0.14
Ours-Fix2.582%0.8752s4.17 ± 0.154.349%0.6956s4.29 ± 0.12
Ours-Normal2.387%0.7091s4.20 ± 0.133.534%0.5283s4.44 ± 0.12

Limitations

The approach requires a two-stage cascaded setup (semantic token generation followed by acoustic decoding via MaskGCT), adding architectural complexity. Evaluation is presently restricted to English and Chinese datasets, leaving multi-lingual scaling beyond these families unverified. Furthermore, the block-size hyperparameter nn requires manual tuning (optimal at n=5n=5), and performance relies on sufficient pre-trained weight initialization from existing large-scale discrete speech models.

Why read this

Speech researchers and engineers working on non-autoregressive text-to-speech should read this paper to learn how to completely decouple generative acoustic models from rigid forced-aligner durations using functional tokens and confidence-guided inference strategies.

Code

Applications

High-fidelity text-to-speech generation, expressive audiobooks, and interactive voice assistants requiring dynamic speaking rates and robust handling of unconstrained text inputs.

Institutions

Tsinghua University, Ant Group

Funding / 經費: National Natural Science Foundation of China, National Social Science Foundation of China, Ant Group

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1069