[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-93331":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":8,"htmlUrl":8,"language":9,"languages":8,"totalLinesOfCode":8,"stars":10,"forks":11,"watchers":12,"openIssues":13,"contributorsCount":13,"subscribersCount":13,"size":13,"stars1d":13,"stars7d":14,"stars30d":14,"stars90d":13,"forks30d":13,"starsTrendScore":15,"compositeScore":16,"rankGlobal":8,"rankLanguage":8,"license":17,"archived":18,"fork":18,"defaultBranch":19,"hasWiki":18,"hasPages":18,"topics":20,"createdAt":8,"pushedAt":8,"updatedAt":21,"readmeContent":22,"aiSummary":23,"trendingCount":13,"starSnapshotCount":13,"syncStatus":24,"lastSyncTime":25,"discoverSource":26},93331,"Wan-Dancer","Wan-Video\u002FWan-Dancer","Wan-Video",null,"Python",324,30,1,0,182,17,71.47,"Apache License 2.0",false,"main",[],"2026-07-22 04:02:08","\n---\n\n# 💃 Wan-Dancer\n\n\u003Cp align=\"center\">\n    💜 \u003Ca href=\"https:\u002F\u002Fhumanaigc.github.io\u002Fwan-dancer-project\u002F\">\u003Cb>Project\u003C\u002Fb>\u003C\u002Fa> &nbsp&nbsp ｜ &nbsp&nbsp 🖥️ \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FWan-Video\u002FWan-Dancer\">GitHub\u003C\u002Fa> &nbsp&nbsp | &nbsp&nbsp🤖 \u003Ca href=\"https:\u002F\u002Fmodelscope.ai\u002Fstudios\u002FWan-AI\u002FWan-Dancer\">MS Space\u003C\u002Fa>&nbsp&nbsp | &nbsp&nbsp🤖 \u003Ca href=\"https:\u002F\u002Fwww.modelscope.cn\u002Fmodels\u002FWan-AI\u002FWan-Dancer-14B\">MS Model\u003C\u002Fa>&nbsp&nbsp | &nbsp&nbsp🤗 \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002FWan-AI\u002FWan-Dancer-14B\">HF Model\u003C\u002Fa>&nbsp&nbsp | &nbsp&nbsp 📑 \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.09581\">Paper\u003C\u002Fa> &nbsp&nbsp \n\u003Cbr>\n\n> **Generating long-duration, high-quality, rhythmic dance videos from music with global structure and temporal continuity**\n\n---\n\n## 📄 Abstract\n\nGenerating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p\u002F30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.\n\n---\n\n## ⚙️ Environment Setup\n\nWe tested the code on:\n\n- **OS**: Ubuntu 22.04\n- **Hardware**: 8 × NVIDIA A800 80GB GPUs\n- **Python**: 3.10.14\n\n### ✅ Install Dependencies\n\n```bash\ncd \u002Fpath\u002Fto\u002FWan-Dancer\npython -m venv venv_wan_dancer\nsource venv_wan_dancer\u002Fbin\u002Factivate\n\n# Install package in editable mode\npip install -e .\n\n# Install additional and specific versions dependencies\npip install moviepy loguru librosa\npip install https:\u002F\u002Fmirrors.aliyun.com\u002Fpytorch-wheels\u002Fcu124\u002Ftorch-2.6.0+cu124-cp310-cp310-linux_x86_64.whl\npip install torchvision==0.21.0\npip install diffusers==0.34.0\npip install yunchang==0.5.0\npip install flash_attn==2.6.3\npip install xfuser==0.4.0\npip install transformers==4.46.2\n```\n\n> 💡 Ensure CUDA 12.4 is installed and compatible with your system.\n\n---\n\n## 🚀 Usage Guide\n\n### 1. 🎬 Generate Global Keyframe Video\n\nRun the global stage script:\n\n```bash\ncd \u002Fpath\u002Fto\u002FWan-Dancer\n.\u002Fgen_video_global.sh\n```\n\n#### 🔧 Important Parameters\n\n| Parameter              | Description |\n|------------------------|-------------|\n| `seed`                 | Random seed for reproducibility. |\n| `image_path`           | Path to reference image. Example: `gen_video\u002Fref_image\u002F1001.jpg` |\n| `prompt_path`          | Path to prompt file (defines dance style).\u003Cbr>Available styles:\u003Cul>\u003Cli>Chinese Classic Dance: `gen_video\u002Fprompt\u002F古典舞_global.txt`\u003C\u002Fli>\u003Cli>K-Pop Dance: `gen_video\u002Fprompt\u002Fkpop_global.txt`\u003C\u002Fli>\u003Cli>Street Dance: `gen_video\u002Fprompt\u002F街舞_global.txt`\u003C\u002Fli>\u003Cli>Tap Dance: `gen_video\u002Fprompt\u002F踢踏舞_global.txt`\u003C\u002Fli>\u003Cli>Latin Dance: `gen_video\u002Fprompt\u002F拉丁舞_global.txt`\u003C\u002Fli>\u003C\u002Ful> |\n| `music_path`           | Path to input music file. Example: `gen_video\u002Fmusic\u002FChineseClassicDance.WAV` |\n| `output_folder`        | Output directory for generated video. |\n| `timestamp`            | Timestamp identifier for output files. |\n| `num_inference_steps`  | Number of diffusion inference steps (e.g., 48). |\n\n\n#### 🌰 Examples\n| Dance Genres | Parameter             | Generated Global Video |\n| ------------ |-----------------------|-----------------|\n| Chinese Classical Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F1001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002F古典舞_global.txt'\u003Cbr>music_path='gen_video\u002Fmusic\u002FChineseClassicDance.WAV'\u003Cbr>num_inference_steps=48\u003Cbr>cfg_scale=5 | [![Chinese Classical Dance](gen_video\u002Fref_image\u002F1001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FmV2fwDpfJ-pODxx6qn-ifq3_UMgbze7P_cI4cLO_vOo.mp4) |\n| Street Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F2001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002F街舞_global.txt'\u003Cbr>music_path='gen_video\u002Fmusic\u002FStreetDance.WAV'\u003Cbr>num_inference_steps=48\u003Cbr>cfg_scale=5 | [![Street Dance](gen_video\u002Fref_image\u002F2001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FMQiVGjY_ngH3imgfIl37xaQoJfbWadYldlZoMWJFMKQ.mp4) |\n| K-Pop Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F3001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002Fkpop_global.txt'\u003Cbr>music_path='gen_video\u002Fmusic_suno\u002F3001.WAV'\u003Cbr>num_inference_steps=48\u003Cbr>cfg_scale=5 | [![K-Pop Dance](gen_video\u002Fref_image\u002F3001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FWGS6Z3VWpgGh8jnt2lrW99XeTB6uu9-H6lCGk1HBLZg.mp4) |\n| Latin Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F4001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002F拉丁舞_global.txt'\u003Cbr>music_path='gen_video\u002Fmusic\u002FLatinDance.WAV'\u003Cbr>num_inference_steps=48\u003Cbr>cfg_scale=5 | [![Latin Dance](gen_video\u002Fref_image\u002F4001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FjnwCUj3WvuErBAxF78b-kttEJoegA6-8VmLMZsayBGI.mp4) |\n| Tap Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F5001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002F踢踏舞_global.txt'\u003Cbr>music_path='gen_video\u002Fmusic\u002FTapDance.wav'\u003Cbr>num_inference_steps=48\u003Cbr>cfg_scale=5 | [![Tap Dance](gen_video\u002Fref_image\u002F5001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FlfrYGNMKzYaLvU3IsMyVJM003T5WZL6QKR7xiifEVAg.mp4)|\n---\n\n### 2. 🎥 Generate Final High-Resolution Video\n\nRun the local refinement stage:\n\n```bash\ncd \u002Fpath\u002Fto\u002FWan-Dancer\n.\u002Fgen_video_local.sh\n```\n\n#### 🔧 Additional Required Parameters\n\n| Parameter             | Description |\n|-----------------------|-------------|\n| `global_video_path`   | Path to the global video generated in Step 2. **Required** for local refinement. |\n| `prompt_path`          | Path to prompt file (defines dance style).\u003Cbr>Available styles:\u003Cul>\u003Cli>Chinese Classic Dance: `gen_video\u002Fprompt\u002F古典舞_local.txt`\u003C\u002Fli>\u003Cli>K-Pop Dance: `gen_video\u002Fprompt\u002Fkpop_local.txt`\u003C\u002Fli>\u003Cli>Street Dance: `gen_video\u002Fprompt\u002F街舞_local.txt`\u003C\u002Fli>\u003Cli>Tap Dance: `gen_video\u002Fprompt\u002F踢踏舞_local.txt`\u003C\u002Fli>\u003Cli>Latin Dance: `gen_video\u002Fprompt\u002F拉丁舞_local.txt`\u003C\u002Fli>\u003C\u002Ful> |\n\n> ✅ All other parameters (`seed`, `image_path`, etc.) are identical to Step 2. \n\n#### 🌰 Examples\n| Dance Genres | Parameter             | Generated Final Video |\n| ------------ |-----------------------|-----------------|\n| Chinese Classical Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F1001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002F古典舞_local.txt'\u003Cbr>music_path='gen_video\u002Fmusic\u002FChineseClassicDance.WAV'\u003Cbr>num_inference_steps=24\u003Cbr>cfg_scale=5\u003Cbr>global_video_path='outputs\u002Fglobal_video\u002F1001_ChineseClassicDance_seed0.mp4' | [![Chinese Classical Dance](gen_video\u002Fref_image\u002F1001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FUycK9FTbYM6imr_6jF9aYbNYTiBggyE0EYptc2TRIAw.mp4) |\n| Street Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F2001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002F街舞_local.txt'\u003Cbr>music_path='gen_video\u002Fmusic\u002FStreetDance.WAV'\u003Cbr>num_inference_steps=24\u003Cbr>cfg_scale=5\u003Cbr>global_video_path='outputs\u002Fglobal_video\u002F2001_StreetDance_seed0.mp4' | [![Street Dance](gen_video\u002Fref_image\u002F2001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FJZtIncJf7zPptZAYsQsoSxA_tyW_r62JfBBikBiTPcY.mp4) |\n| K-Pop Dance | seed=100\u003Cbr>image_path='gen_video\u002Fref_image\u002F3001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002Fkpop_local.txt'\u003Cbr>music_path='gen_video\u002Fmusic_suno\u002F3001.WAV'\u003Cbr>num_inference_steps=24\u003Cbr>cfg_scale=5\u003Cbr>global_video_path='outputs\u002Fglobal_video\u002F3001_KPopDance_seed0.mp4' | [![K-Pop Dance](gen_video\u002Fref_image\u002F3001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FSi5ze8sR0Rm-aPUGSKsTJ2PXJAu3HtnVAzEPM85bkrc.mp4) |\n| Latin Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F4001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002F拉丁舞_local.txt'\u003Cbr>music_path='gen_video\u002Fmusic\u002FLatinDance.WAV'\u003Cbr>num_inference_steps=24\u003Cbr>cfg_scale=5\u003Cbr>global_video_path='outputs\u002Fglobal_video\u002F4001_LatinDance_seed0.mp4' | [![Latin Dance](gen_video\u002Fref_image\u002F4001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FkL-0AAqQtigvaidF8Xa8YeTIs4pDLOa_4n5nqXmYiRk.mp4) |\n| Tap Dance | seed=0\u003Cbr>image_path='gen_video\u002Fref_image\u002F5001.jpg'\u003Cbr>prompt_path='gen_video\u002Fprompt\u002F踢踏舞_local.txt'\u003Cbr>music_path='gen_video\u002Fmusic\u002FTapDance.wav'\u003Cbr>num_inference_steps=24\u003Cbr>cfg_scale=5\u003Cbr>global_video_path='outputs\u002Fglobal_video\u002F5001_TapDance_seed0.mp4' | [![Tap Dance](gen_video\u002Fref_image\u002F5001.jpg)](https:\u002F\u002Fcloud.video.taobao.com\u002Fvod\u002FGbnX-XzekrvNulbbDMw_2kEotadZmUT6KFY5smTkNZ0.mp4) |\n\n\u003Cstrong>Note:\u003C\u002Fstrong> The `num_inference_steps` should be set to a larger value (e.g., 48) for longer time videos.\n\n---\n\n## 🙏 Acknowledgements\n\nThis work builds upon and integrates components from the following open-source projects:\n\n1. [DiffSynth-Studio](https:\u002F\u002Fgithub.com\u002Fmodelscope\u002FDiffSynth-Studio)  \n2. [Wan2.1](https:\u002F\u002Fgithub.com\u002FWan-Video\u002FWan2.1)  \n\n---\n\n## 📜 License\n\nThis project is licensed under the Apache 2.0 License — see the [LICENSE](LICENSE) file for details.\n\n## 📚 Citation\n\nIf you use this code or framework in your research, please cite:\n\n```bibtex\n@article{wan-dancer-2026,\n  title={Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation},\n  author={Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Bang Zhang},\n  website={https:\u002F\u002Fhumanaigc.github.io\u002Fwan-dancer-project\u002F},\n  url={https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.09581},\n  year={2026}\n}\n\n","Wan-Dancer 是一个面向音乐驱动长时序舞蹈视频生成的开源框架，专注于解决扩散模型在生成超过20秒高质量舞蹈视频时存在的时序漂移、身份不一致和动作重复等问题。其核心采用分层架构：先基于全曲音乐进行全局关键帧规划，再通过局部时序精修实现高保真运动合成；关键技术包括时间映射RoPE嵌入实现动态帧率对齐、光流引导损失增强运动连续性、以及运动速度可控机制保障快速动作细节。适用于舞蹈内容创作、AI虚拟偶像驱动、数字人表演生成等需分钟级（>60秒）、720p\u002F30fps、节奏严格同步的音乐-舞蹈协同生成场景。",2,"2026-07-16 02:30:07","CREATED_QUERY"]