---
id: im26_interspeech
title: "PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation"
authors:
  - Jaekwon Im
  - Natalia Polouliakh
  - Taketo Akama
year: 2026
doi: 10.21437/Interspeech.2026-248
isca_url: https://www.isca-archive.org/interspeech_2026/im26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/im26_interspeech.pdf
session: Streaming Speech Synthesis
topics:
  - speech-llm
  - self-supervised
  - dataset
category: audio-understanding
labels:
  - generative-model
institutions:
  - KAIST
  - Sony Computer Science Laboratories
code:
  url: https://jakeoneijk.github.io/pfd2m_project
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: im26_interspeech
  category: audio-understanding
  labels:
    - generative-model
  institutions:
    - KAIST
    - Sony Computer Science Laboratories
  code: https://jakeoneijk.github.io/pfd2m_project
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-248
  pdf: https://www.isca-archive.org/interspeech_2026/im26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/im26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/im26_interspeech/markdown.md
---

# PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation

*Jaekwon Im, Natalia Polouliakh, Taketo Akama*

[PDF](https://www.isca-archive.org/interspeech_2026/im26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/im26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-248)

**Category:** `audio-understanding` · **Labels:** `generative-model`

**TL;DR** — PF-D2M is a pose-free diffusion model for universal dance-to-music generation that uses video visual features and a progressive multi-stage training recipe, achieving state-of-the-art performance on rhythm alignment and music quality.

## Key contributions

- Proposes a pose-free universal dance-to-music generation framework (PF-D2M) capable of handling multiple and non-human dancers without brittle keypoint extractors.
- Incorporates a pre-trained Synchformer visual encoder to extract rich temporal visual features directly from raw dance frames.
- Introduces a progressive three-stage training strategy (Text-to-Audio initialization, VGGSound alignment, and multimodal fine-tuning) to combat severe data scarcity.
- Curates a filtered 191-hour text-to-music and video-to-audio training diet using automated source separation and Voice Activity Detection.

## Problem

Prior dance-to-music generation models rely heavily on motion features extracted from a single human dancer via 3D SMPL models or 2D keypoints (e.g., HRNet), making them fail on multiple performers, animated characters, or complex in-the-wild camera angles. Furthermore, existing public datasets like AIST++ provide only 60 unique songs, causing extreme overfitting and memorization in deep generative models. Overcoming this data bottleneck is critical for producing musically diverse, studio-quality, synchronized audio for real-world creative workflows.

## Method

PF-D2M adopts the Stable Audio Open architecture utilizing a DiT backbone and a pre-trained VAE that compresses 44.1 kHz stereo audio into latent representations. The model conditions on three sources: text captions via a T5-base cross-attention mechanism, diffusion timesteps via sinusoidal embeddings, and dense visual features extracted from video frames (at 25 fps) via Synchformer.

The visual features are upsampled via nearest-neighbor interpolation, projected using 1D convolutions to match channel dimensions, and concatenated directly with the DiT input. Simultaneously, they are projected via linear layers and injected into every DiT layer using frame-wise scales and biases in adaptive layer normalization (AdaLN) alongside timestep embeddings. Both text and visual conditioning are dropped out with a 10% probability for classifier-free guidance.

The training recipe is split into three progressive stages: Stage 0 initializes weights from Stable Audio Open while zero-initializing new modules. Stage 1 trains on 500 hours of VGGSound for general audio-visual synchronization using Qwen-Audio captions. Stage 2 fine-tunes on a multi-modal mixture of AIST++ (dance-to-music), filtered FMA and MoisesDB (text-to-music), and VGGSound in a 2:4:1 dataset ratio. Text prompts are constructed stochastically from tags generated by Qwen2-Audio. Inference employs DPM-Solver++ with 100 steps and a guidance scale of 5.0.

## Experimental setup

Evaluated on the AIST++ test set (reserving specific unseen tracks mBR0, mMH0, mLO2, and mJB5) and an in-the-wild benchmark of 20 challenging videos spanning single/multiple human and non-human dancers. Compared against CDCD, LORIS, and Text-Inv using objective rhythm metrics (BCS, CSD, BHS, HSD, F1) and subjective 5-point Likert scale listening tests administered to 20 participants. Implemented with batch size 128 using AdamW optimizer.

## Results

On the AIST++ test set, PF-D2M (Stage 2) achieves state-of-the-art results across key rhythm metrics, yielding a Beat Hit Score (BHS) of 99.8% (vs 95.3% for LORIS and 80.9% for Text-Inv) and an F1 score of 94.3%. In subjective evaluations across in-the-wild categories, PF-D2M significantly outperforms baselines in both perceptual music quality and dance-music alignment, showing exceptional robustness on multi-cut and multi-dancer videos.

| Method | BCS↑ | CSD↓ | BHS↑ | HSD↓ | F1↑ |
|---|---|---|---|---|---|
| CDCD | 89.2 | 9.0 | 93.8 | 10.0 | 91.5 |
| LORIS | 89.9 | 8.9 | 95.3 | 8.9 | 92.5 |
| Textual-Inv | 90.6 | 11.1 | 80.9 | 28.3 | 85.5 |
| PF-D2M (S1) | 90.5 | 13.1 | 91.2 | 18.6 | 90.9 |
| PF-D2M (S2) | 89.4 | 8.1 | 99.8 | 1.9 | 94.3 |

## Limitations

The generated audio clips are restricted to relatively short durations (7.98-second training clips, 5.12-second evaluation windows) due to underlying diffusion architecture constraints, preventing the generation of full-length, structurally progressive musical tracks. The lack of standardized, large-scale dance-to-music evaluation benchmarks also limits comprehensive objective validation.

## Why read this

Researchers and audio-video engineers should read this paper to see how replacing brittle pose estimation with raw-video visual transformers (Synchformer) and multi-stage progressive training dramatically expands the domain generality of audio generation models.

## Code

- https://jakeoneijk.github.io/pfd2m_project

## Applications

Automated choreography sound-tracking, video content creation tools, and real-time interactive performance systems for digital avatars and human dancers.

## Institutions / 機構

KAIST, Sony Computer Science Laboratories

## Related

- [GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment](wang26da_interspeech.md) — same problem · relatedness 3.0/3
- [FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision](wang26b_interspeech.md) — shared technique · relatedness 2.1/3
- [ARCHES: An Agent-Based Refinement Cycle for Hierarchical Synthesis of Sound Effects for Variety Shows](lei26_interspeech.md) — relatedness 1.8/3
- [Listening to Motion in Space: Vision-Grounded Event-wise Video-to-Audio Generation and Rendering](park26m_interspeech.md) — same problem · relatedness 1.8/3
- [FoleyImmersive: Decoupling What and Where for Video-to-First-Order Ambisonics](liang26b_interspeech.md) — relatedness 1.7/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
