All papers
Audio understandingFull-paper digest

PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation

Jaekwon Im, Natalia Polouliakh, Taketo Akama

Code & resourcesjakeoneijk.github.io/pfd2m_project

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.0 KB · Ready to paste

Preview copied content

TL;DR — PF-D2M is a pose-free diffusion model for universal dance-to-music generation that uses video visual features and a progressive multi-stage training recipe, achieving state-of-the-art performance on rhythm alignment and music quality.

Key contributions

  • Proposes a pose-free universal dance-to-music generation framework (PF-D2M) capable of handling multiple and non-human dancers without brittle keypoint extractors.
  • Incorporates a pre-trained Synchformer visual encoder to extract rich temporal visual features directly from raw dance frames.
  • Introduces a progressive three-stage training strategy (Text-to-Audio initialization, VGGSound alignment, and multimodal fine-tuning) to combat severe data scarcity.
  • Curates a filtered 191-hour text-to-music and video-to-audio training diet using automated source separation and Voice Activity Detection.

Problem

Prior dance-to-music generation models rely heavily on motion features extracted from a single human dancer via 3D SMPL models or 2D keypoints (e.g., HRNet), making them fail on multiple performers, animated characters, or complex in-the-wild camera angles. Furthermore, existing public datasets like AIST++ provide only 60 unique songs, causing extreme overfitting and memorization in deep generative models. Overcoming this data bottleneck is critical for producing musically diverse, studio-quality, synchronized audio for real-world creative workflows.

Method

PF-D2M adopts the Stable Audio Open architecture utilizing a DiT backbone and a pre-trained VAE that compresses 44.1 kHz stereo audio into latent representations. The model conditions on three sources: text captions via a T5-base cross-attention mechanism, diffusion timesteps via sinusoidal embeddings, and dense visual features extracted from video frames (at 25 fps) via Synchformer.

The visual features are upsampled via nearest-neighbor interpolation, projected using 1D convolutions to match channel dimensions, and concatenated directly with the DiT input. Simultaneously, they are projected via linear layers and injected into every DiT layer using frame-wise scales and biases in adaptive layer normalization (AdaLN) alongside timestep embeddings. Both text and visual conditioning are dropped out with a 10% probability for classifier-free guidance.

The training recipe is split into three progressive stages: Stage 0 initializes weights from Stable Audio Open while zero-initializing new modules. Stage 1 trains on 500 hours of VGGSound for general audio-visual synchronization using Qwen-Audio captions. Stage 2 fine-tunes on a multi-modal mixture of AIST++ (dance-to-music), filtered FMA and MoisesDB (text-to-music), and VGGSound in a 2:4:1 dataset ratio. Text prompts are constructed stochastically from tags generated by Qwen2-Audio. Inference employs DPM-Solver++ with 100 steps and a guidance scale of 5.0.

Experimental setup

Evaluated on the AIST++ test set (reserving specific unseen tracks mBR0, mMH0, mLO2, and mJB5) and an in-the-wild benchmark of 20 challenging videos spanning single/multiple human and non-human dancers. Compared against CDCD, LORIS, and Text-Inv using objective rhythm metrics (BCS, CSD, BHS, HSD, F1) and subjective 5-point Likert scale listening tests administered to 20 participants. Implemented with batch size 128 using AdamW optimizer.

Results

On the AIST++ test set, PF-D2M (Stage 2) achieves state-of-the-art results across key rhythm metrics, yielding a Beat Hit Score (BHS) of 99.8% (vs 95.3% for LORIS and 80.9% for Text-Inv) and an F1 score of 94.3%. In subjective evaluations across in-the-wild categories, PF-D2M significantly outperforms baselines in both perceptual music quality and dance-music alignment, showing exceptional robustness on multi-cut and multi-dancer videos.

MethodBCS↑CSD↓BHS↑HSD↓F1↑
CDCD89.29.093.810.091.5
LORIS89.98.995.38.992.5
Textual-Inv90.611.180.928.385.5
PF-D2M (S1)90.513.191.218.690.9
PF-D2M (S2)89.48.199.81.994.3

Limitations

The generated audio clips are restricted to relatively short durations (7.98-second training clips, 5.12-second evaluation windows) due to underlying diffusion architecture constraints, preventing the generation of full-length, structurally progressive musical tracks. The lack of standardized, large-scale dance-to-music evaluation benchmarks also limits comprehensive objective validation.

Why read this

Researchers and audio-video engineers should read this paper to see how replacing brittle pose estimation with raw-video visual transformers (Synchformer) and multi-stage progressive training dramatically expands the domain generality of audio generation models.

Code

Applications

Automated choreography sound-tracking, video content creation tools, and real-time interactive performance systems for digital avatars and human dancers.

Institutions

KAIST, Sony Computer Science Laboratories

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-248