Wan Dancer

Turn a single portrait and a music track into a beat-synced dance video at 720p and 30fps — coherent well past the 20-second wall that stops most models. Wan Dancer (Wan-Dancer-14B) is Alibaba Tongyi Lab's open-source music-to-dance model, released under Apache-2.0.

Hosted generator coming soon

720p

30fps video output

1 min+

Coherent past the 20s wall

5

Dance genres supported

Apache-2.0

Open weights, self-hostable

What Is Wan Dancer

A music-driven portrait dance video model from Alibaba Tongyi Lab.

Wan-Dancer-14B generates rhythm-synced dance video from a single reference photo and a music track — no motion capture and no reference dance video required.

Wan Dancer, released as Wan-Dancer-14B by Alibaba's Tongyi Lab, is a music-to-dance video model. You start from one reference photo of a person and a music track, and Wan Dancer choreographs and renders that person dancing in time with the music at 720p and 30fps. Where most video models drift or fall apart after about 20 seconds, Wan Dancer stays coherent for over a minute.

It works by planning before it renders. A global stage reads the full music track and lays out the long-range dance as keyframes, then a local stage refines the motion between them frame by frame. That two-stage design, built on the Wan image-to-video backbone, is what lets you go from a still image and a song to a full dance clip instead of a few seconds of movement. The hosted generator is coming soon — until then, this page covers exactly how the model works and where it fits.

Best for

Music videos and lyric clips, dance-cover content for TikTok and Reels, virtual performers and VTuber intros, branded promo loops, and anyone who wants a photo to dance to a specific song.

How It Works

From photo and music to dance in three stages.

Wan Dancer separates planning from rendering — the generator isn't live yet, but this is the workflow the model follows.

01

Give it a portrait and a track

Provide one clear reference image of a person — a vertical, full-body shot works best — plus a music or audio file and a short text prompt naming the dance style. Clean, low-noise audio and a clearly defined subject give the steadiest results.

02

The global stage plans the choreography

Wan Dancer reads the entire music track and lays out keyframes for the whole routine, setting the long-range structure so the dance holds together from start to finish. Planning from the full track is what buys coherence beyond 20 seconds.

03

The local stage renders the motion

A second stage refines the frames between keyframes into smooth, beat-aligned motion at 720p and 30fps, then outputs the finished dance video as MP4. Motion-speed control keeps detail sharp even during fast passages.

Key Features

Audio-driven choreography that keeps your subject intact.

A music-to-dance video generator built for long, rhythm-locked, single-subject choreography.

Music-driven dance generation

Wan Dancer generates the dance directly from your audio, so choreography follows the actual music instead of a generic loop. You supply the track up front rather than syncing music to a silent clip afterward. A text prompt guides style and mood on top of the audio.

Minute-scale temporal coherence

The global-to-local pipeline holds structure past the roughly 20-second point where most diffusion models fall apart, producing dances that run over a minute — long enough for a full chorus or hook.

Single-photo identity lock

One portrait is enough. Wan Dancer carries the reference person's face, hair, and outfit through the whole routine, so the dancer stays recognizable from the first frame to the last.

720p at 30fps output

Wan Dancer renders high-definition video at a smooth 30 frames per second, ready for short-form social clips and music-video cuts on TikTok, Reels, and Shorts.

Five dance genres

Trained across Chinese classical, K-pop, street, tap, and Latin, Wan Dancer can match a range of musical styles from a single reference image — you pick the style through your prompt.

Open weights, ComfyUI-ready

Wan-Dancer-14B is released under Apache-2.0 on Hugging Face and ModelScope, with inference code, ComfyUI integration, and LoRA fine-tuning for custom choreography. Self-host it and use outputs commercially.

Use Cases

Where a photo-to-dance model earns its place.

Turn a single portrait and a track into dance video for the places short-form lives.

Music & lyric videos

Generate dance sequences that follow a track for lyric videos, visualizers, and music-video inserts — useful for indie musicians and promo teasers without booking a shoot.

TikTok & Reels dance clips

Turn one portrait into a dance short synced to a trending song, then batch variations for TikTok, Reels, and Shorts with real beat sync.

Virtual performers & VTubers

Give a virtual influencer or VTuber avatar a dancing intro, stinger, or loop that reacts to music — no rigging or motion-capture pipeline needed.

Branded dance spots

Animate a brand character or spokesperson dancing to campaign audio for social ads, landing pages, and product launches.

What Creators Say

Why Wan Dancer earns a spot in the workflow.

Trustpilot

Trusted by music, dance, and short-form creators — open-source under Apache-2.0.

Lena Marsh, Music Video Editor

Lena Marsh

Music Video Editor

"The music-first approach changed my workflow. I feed in the track and get choreography that actually fits it, instead of cutting a silent clip to the beat later."

Amara Sy, Dance Content Producer

Amara Sy

Dance Content Producer

"Five genres from one reference image means I can match Latin one day and street the next without re-shooting anything."

Nadia Hult, Motion Designer

Nadia Hult

Motion Designer

"Beat alignment holds up across the whole clip. The motion lands on the rhythm, which is exactly what silent-then-synced tools never nailed."

Yara Kwon, Virtual Influencer Manager

Yara Kwon

Virtual Influencer Manager

"For virtual-influencer content, identity preservation matters. The dancer stays recognizably the same character from start to finish."

Diego Ramos, Short-Form Creator

Diego Ramos

Short-Form Creator

"Getting past the 20-second wall is the whole thing for me. A dance that stays coherent for a full minute is finally usable in a real video."

Jonas Pearl, Pipeline TD

Jonas Pearl

Pipeline TD

"Open weights and ComfyUI support sold our team. We fine-tune choreography with LoRA and run everything in our own pipeline."

Explore more AI video and image models

FAQ

Wan Dancer FAQ

What is Wan Dancer?

Wan Dancer (Wan-Dancer-14B) is an open-source music-to-dance video model from Alibaba's Tongyi Lab. It takes one portrait photo and a music track and generates a video of that person dancing in time with the music at 720p and 30fps.

How does Wan Dancer generate a dance video?

It uses a two-stage pipeline. A global stage reads the full music track and plans the dance as keyframes, then a local stage refines the motion between them. This planning-first design is what keeps long dances coherent.

How long can a Wan Dancer video be?

Wan Dancer is built for minute-scale generation and stays coherent well past the roughly 20-second point where most diffusion video models break down. Project demos run over two minutes; the open-source build targets shorter music inputs.

What do I need to use Wan Dancer?

A single reference image of a person (a vertical, full-body shot works best), a music or audio file, and a short text prompt for the dance style. Clean, low-noise audio produces the steadiest results.

What dance styles does Wan Dancer support?

It was trained across five genres: Chinese classical, K-pop, street, tap, and Latin. You choose the style through your text prompt alongside the music.

Is Wan Dancer open source?

Yes. Wan-Dancer-14B is released under the Apache-2.0 license on Hugging Face and ModelScope, with inference code, ComfyUI integration, and LoRA fine-tuning support.

Can I use Wan Dancer on this site right now?

A hosted Wan Dancer generator on wan27.org is coming soon. In the meantime, this page explains how the model works, and the open weights are available on Hugging Face and ModelScope.

How is Wan Dancer different from Wan Animate?

Wan Dancer is driven by music — it choreographs a dance from an audio track. Wan Animate is driven by a reference video and handles character animation and replacement. They are sibling models from the same lab but solve different tasks.