All papers
Speech synthesisFull-paper digest

DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching

Yuepeng Jiang, Huakang Chen, Ziqian Ning, Jixun Yao, zerui Han, Di Wu, Meng Meng, Jian Luan, Zhonghua Fu, Lei Xie

Code & resourcesgithub.com/xiaomi-research/diffrhythm2

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.5 KB · Ready to paste

Preview copied content

TL;DR — DiffRhythm 2 is an end-to-end semi-autoregressive song generation framework utilizing block flow matching, a 5 Hz music VAE, and cross-pair preference optimization to produce 210-second songs with faithful lyric alignment and strong open-source benchmark performance.

Key contributions

  • A semi-autoregressive block flow matching architecture that achieves reliable lyric-vocal alignment without external duration labels or explicit constraints.
  • Stochastic block representation alignment (REPA) loss using MuQ representations to improve long-sequence structural coherence and musicality.
  • Cross-pair preference optimization (CPPO) that groups and jointly optimizes conflicting/synergistic preference pairs in a single model, avoiding weight-merging degradation.
  • A customized 5 Hz music VAE that achieves a 4800× encoding and 9600× decoding compression ratio to enable feasible long-sequence modeling.

Problem

Generating full-length songs exceeding three minutes requires joint modeling of lyrics, structure, singing vocals, and accompaniment. Existing non-autoregressive models struggle with long-sequence lyric-vocal alignment unless constrained by sentence-level timestamps or intermediate representations, which often hurt creativity and musicality. Meanwhile, autoregressive alternatives suffer from slow inference speeds, and multi-preference alignment via independent DPO models followed by weight interpolation causes performance degradation due to averaging effects.

Method

DiffRhythm 2 combines a 5 Hz music VAE with a Diffusion Transformer. The VAE processes 24 kHz audio and reconstructs it at 48 kHz, employing Stable Audio 2 VAE encoder architecture, an intermediate transformer block, and a BigVGAN decoder optimized via multi-scale mel, multi-scale STFT, and CQT/multi-period/multi-scale discriminators.

The generative backbone uses block flow matching. A target latent sequence ZZ of length ll is split into blocks of size bb (k=⌈l/b⌉k = \lceil l/b \rceil). Each block is generated via flow matching conditioning on style prompts SS, lyrics LL, timestep tt, and preceding blocks via a block-level causal attention mask. The input format stacks clean and noisy block sequences paired with an attention mask where the ii-th block attends to clean blocks 11 through i−1i-1 and its own noisy block. Timesteps are fixed to −1-1 for prompts/lyrics, 11 for clean sequences, and sampled from U[0,1]U[0, 1] independently for noisy blocks.

Variable length generation is supported by appending nn End-of-Prediction (EOP) frames modeled as a constant vector of ones (N(1,0)N(1, 0)) to the final block. To guide structure, a stochastic block REPA loss randomly samples 10 blocks per noisy sequence against MuQ target representations. For multi-preference tuning, cross-pair preference optimization pairs (musicality, lyric alignment) and (style similarity, audio quality) during DPO training, where winning samples satisfy both preferences and losing samples satisfy at least one.

Experimental setup

Trained on 1.4 million songs comprising roughly 70,000 hours of Chinese, English, and instrumental music (4:5:1 ratio). Evaluated on a test set of 50 real and 50 generated lyrics paired with 3 randomized style prompts each (300 total cases). Compared against commercial baselines (Suno V4.5, Mureka-O1) and open-source models (DiffRhythm+, ACE-Step, LeVo). Evaluated via professional human MOS (MUS, HAR, VOC, ACC, OVP) and objective metrics (PER, Mulan-T, Mulan-A, Audiobox-Aesthetics, SongEval).

Results

DiffRhythm 2 achieves a Phoneme Error Rate (PER) of 0.13 and Mulan-T text prompt style similarity of 0.40, outperforming open-source baselines. In subjective evaluations, it attains a musicality (MUS) score of 3.57, vocal-accompaniment harmony (HAR) of 3.81, and overall performance (OVP) of 3.77, beating ACE-Step (OVP 3.55) and LeVo (OVP 3.56). Ablation studies show that removing CPPO degrades PER to 0.18 and lowers SongEval overall musicality (MU) from 3.93 to 3.57, while removing REPA loss leads to structural alignment failure.

ModelPER ↓\downarrowMulan-T ↑\uparrowMUS ↑\uparrowHAR ↑\uparrowOVP ↑\uparrow
SUNO V4.50.280.383.684.033.92
Mureka-O10.090.373.713.993.87
DiffRhythm+0.150.253.103.223.27
ACE-Step0.230.283.403.753.55
LeVo0.190.353.483.683.56
DiffRhythm 20.130.403.573.813.77

Limitations

The 5 Hz low-frame-rate VAE compresses audio heavily, imposing an upper bound on reconstructed audio fidelity compared to real recordings. Global modeling of audio style prompts limits fine-grained stylistic capture (lower Mulan-A score than LeVo). Open-source models still lag behind commercial benchmarks like Suno and Mureka in overall vocal quality and emotional nuance.

Why read this

Read this if you want to understand how block flow matching and cross-pair preference optimization can solve long-sequence alignment and reward-merging degradation in generative audio without autoregressive latency bottlenecks.

Code

Applications

Full-length controllable song generation, interactive music creation platforms, and multi-track audio generation.

Institutions

Northwestern Polytechnical University, Xiaomi

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-128