720p
30fps video output
Turn a single portrait and a music track into a beat-synced dance video at 720p and 30fps — coherent well past the 20-second wall that stops most models. Wan Dancer (Wan-Dancer-14B) is Alibaba Tongyi Lab's open-source music-to-dance model, released under Apache-2.0.
Hosted generator coming soon
720p
30fps video output
1 min+
Coherent past the 20s wall
5
Dance genres supported
Apache-2.0
Open weights, self-hostable
What Is Wan Dancer
Wan-Dancer-14B generates rhythm-synced dance video from a single reference photo and a music track — no motion capture and no reference dance video required.
Wan Dancer, released as Wan-Dancer-14B by Alibaba's Tongyi Lab, is a music-to-dance video model. You start from one reference photo of a person and a music track, and Wan Dancer choreographs and renders that person dancing in time with the music at 720p and 30fps. Where most video models drift or fall apart after about 20 seconds, Wan Dancer stays coherent for over a minute.
It works by planning before it renders. A global stage reads the full music track and lays out the long-range dance as keyframes, then a local stage refines the motion between them frame by frame. That two-stage design, built on the Wan image-to-video backbone, is what lets you go from a still image and a song to a full dance clip instead of a few seconds of movement. The hosted generator is coming soon — until then, this page covers exactly how the model works and where it fits.
Best for
Music videos and lyric clips, dance-cover content for TikTok and Reels, virtual performers and VTuber intros, branded promo loops, and anyone who wants a photo to dance to a specific song.
How It Works
Wan Dancer separates planning from rendering — the generator isn't live yet, but this is the workflow the model follows.
Provide one clear reference image of a person — a vertical, full-body shot works best — plus a music or audio file and a short text prompt naming the dance style. Clean, low-noise audio and a clearly defined subject give the steadiest results.
Wan Dancer reads the entire music track and lays out keyframes for the whole routine, setting the long-range structure so the dance holds together from start to finish. Planning from the full track is what buys coherence beyond 20 seconds.
A second stage refines the frames between keyframes into smooth, beat-aligned motion at 720p and 30fps, then outputs the finished dance video as MP4. Motion-speed control keeps detail sharp even during fast passages.
Key Features
A music-to-dance video generator built for long, rhythm-locked, single-subject choreography.
Wan Dancer generates the dance directly from your audio, so choreography follows the actual music instead of a generic loop. You supply the track up front rather than syncing music to a silent clip afterward. A text prompt guides style and mood on top of the audio.
The global-to-local pipeline holds structure past the roughly 20-second point where most diffusion models fall apart, producing dances that run over a minute — long enough for a full chorus or hook.
One portrait is enough. Wan Dancer carries the reference person's face, hair, and outfit through the whole routine, so the dancer stays recognizable from the first frame to the last.
Wan Dancer renders high-definition video at a smooth 30 frames per second, ready for short-form social clips and music-video cuts on TikTok, Reels, and Shorts.
Trained across Chinese classical, K-pop, street, tap, and Latin, Wan Dancer can match a range of musical styles from a single reference image — you pick the style through your prompt.
Wan-Dancer-14B is released under Apache-2.0 on Hugging Face and ModelScope, with inference code, ComfyUI integration, and LoRA fine-tuning for custom choreography. Self-host it and use outputs commercially.
Use Cases
Turn a single portrait and a track into dance video for the places short-form lives.
Generate dance sequences that follow a track for lyric videos, visualizers, and music-video inserts — useful for indie musicians and promo teasers without booking a shoot.
Turn one portrait into a dance short synced to a trending song, then batch variations for TikTok, Reels, and Shorts with real beat sync.
Give a virtual influencer or VTuber avatar a dancing intro, stinger, or loop that reacts to music — no rigging or motion-capture pipeline needed.
Animate a brand character or spokesperson dancing to campaign audio for social ads, landing pages, and product launches.
What Creators Say
Trusted by music, dance, and short-form creators — open-source under Apache-2.0.

Lena Marsh
Music Video Editor
"The music-first approach changed my workflow. I feed in the track and get choreography that actually fits it, instead of cutting a silent clip to the beat later."

Amara Sy
Dance Content Producer
"Five genres from one reference image means I can match Latin one day and street the next without re-shooting anything."

Nadia Hult
Motion Designer
"Beat alignment holds up across the whole clip. The motion lands on the rhythm, which is exactly what silent-then-synced tools never nailed."

Yara Kwon
Virtual Influencer Manager
"For virtual-influencer content, identity preservation matters. The dancer stays recognizably the same character from start to finish."

Diego Ramos
Short-Form Creator
"Getting past the 20-second wall is the whole thing for me. A dance that stays coherent for a full minute is finally usable in a real video."

Jonas Pearl
Pipeline TD
"Open weights and ComfyUI support sold our team. We fine-tune choreography with LoRA and run everything in our own pipeline."

Generate and edit images with precise visual control.

Text to video, image to video, and reference video workflows.

Main overview for the Wan 2.5 model family.

Explore text-to-video, image-to-video, speech, and animate workflows.

Create image assets with strong typography and layout control.

Plan image-to-video prompts with motion, camera, and sound.
FAQ
Wan Dancer (Wan-Dancer-14B) is an open-source music-to-dance video model from Alibaba's Tongyi Lab. It takes one portrait photo and a music track and generates a video of that person dancing in time with the music at 720p and 30fps.
It uses a two-stage pipeline. A global stage reads the full music track and plans the dance as keyframes, then a local stage refines the motion between them. This planning-first design is what keeps long dances coherent.
Wan Dancer is built for minute-scale generation and stays coherent well past the roughly 20-second point where most diffusion video models break down. Project demos run over two minutes; the open-source build targets shorter music inputs.
A single reference image of a person (a vertical, full-body shot works best), a music or audio file, and a short text prompt for the dance style. Clean, low-noise audio produces the steadiest results.
It was trained across five genres: Chinese classical, K-pop, street, tap, and Latin. You choose the style through your text prompt alongside the music.
Yes. Wan-Dancer-14B is released under the Apache-2.0 license on Hugging Face and ModelScope, with inference code, ComfyUI integration, and LoRA fine-tuning support.
A hosted Wan Dancer generator on wan27.org is coming soon. In the meantime, this page explains how the model works, and the open weights are available on Hugging Face and ModelScope.
Wan Dancer is driven by music — it choreographs a dance from an audio track. Wan Animate is driven by a reference video and handles character animation and replacement. They are sibling models from the same lab but solve different tasks.