[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-96151":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":8,"htmlUrl":8,"language":9,"languages":8,"totalLinesOfCode":8,"stars":10,"forks":11,"watchers":12,"openIssues":13,"contributorsCount":14,"subscribersCount":14,"size":14,"stars1d":14,"stars7d":14,"stars30d":15,"stars90d":14,"forks30d":14,"starsTrendScore":14,"compositeScore":16,"rankGlobal":8,"rankLanguage":8,"license":17,"archived":18,"fork":18,"defaultBranch":19,"hasWiki":18,"hasPages":18,"topics":20,"createdAt":8,"pushedAt":8,"updatedAt":21,"readmeContent":22,"aiSummary":23,"trendingCount":14,"starSnapshotCount":14,"syncStatus":24,"lastSyncTime":25,"discoverSource":26},96151,"ComfyUI-Ref2VA-VSA","Kablex\u002FComfyUI-Ref2VA-VSA","Kablex",null,"Python",232,28,103,1,0,42,45.59,"Other",false,"main",[],"2026-09-20 04:01:32","# ComfyUI-Ref2VA-VSA: Ultra-Fast Character Video Generation (72s on RTX 4090)\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FComfyUI-Custom%20Node-blue?style=for-the-badge\" alt=\"ComfyUI\">\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FModel-MiniMax--H3%20Ref2VA-orange?style=for-the-badge\" alt=\"MiniMax-H3\">\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FSpeed-72s%20%2F%205s%20Video-brightgreen?style=for-the-badge\" alt=\"Speed\">\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FHardware-RTX%204090%20(24GB)-purple?style=for-the-badge\" alt=\"Hardware\">\n\u003C\u002Fp>\n\n> **Breakthrough inference speed for MiniMax H3 Reference-to-Video (Ref2VA \u002F R2VA)**: Generate a full 5-second, 24 fps video conditioned on reference character images in **~72s (warm) \u002F 95s (first run)** on a single consumer **NVIDIA GeForce RTX 4090 (24GB)**.\n\n---\n\n## 🎬 Empirical Benchmark & Results\n\nBoth videos were generated on the exact same **NVIDIA GeForce RTX 4090 (24GB)**, using the exact same prompt, character reference image, seed (`981445682258077`), and resolution (**1344×768, 5.16s @ 24 fps \u002F 124 frames**).\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\"assets\u002Fexample_character.jpg\" width=\"220\" style=\"border-radius: 8px;\" alt=\"Input Reference Character\">\n  \u003Cbr>\n  \u003Cb>Input Character Reference Image\u003C\u002Fb>\n\u003C\u002Fp>\n\n\u003Ctable>\n  \u003Cthead>\n    \u003Ctr align=\"center\">\n      \u003Cth width=\"50%\">\n        \u003Ch3>🚀 Ref2VA VSA (Ours - 4 Steps)\u003C\u002Fh3>\n        \u003Cb>⏱️ 95s (1st run) \u002F ~72s (warm)\u003C\u002Fb>\n      \u003C\u002Fth>\n      \u003Cth width=\"50%\">\n        \u003Ch3>🔬 Video Delta Net \u002F VDN-H3 (8 Steps)\u003C\u002Fh3>\n        \u003Cb>⏱️ 213s (1st run) \u002F ~135s (warm)\u003C\u002Fb>\n      \u003C\u002Fth>\n    \u003C\u002Ftr>\n  \u003C\u002Fthead>\n  \u003Ctbody>\n    \u003Ctr>\n      \u003Ctd align=\"center\">\n        \u003Cimg src=\"assets\u002Fvsa_ref2va_preview.webp\" width=\"100%\" alt=\"Ref2VA VSA Preview\">\n        \u003Cbr>\n        \u003Ca href=\"assets\u002Fvsa_ref2va_4step_95s.mp4\">\u003Cb>▶ Download Full 1344x768 Video (MP4)\u003C\u002Fb>\u003C\u002Fa>\n      \u003C\u002Ftd>\n      \u003Ctd align=\"center\">\n        \u003Cimg src=\"assets\u002Fvdn_ref2va_preview.webp\" width=\"100%\" alt=\"VDN-H3 Preview\">\n        \u003Cbr>\n        \u003Ca href=\"assets\u002Fvdn_ref2va_8step_213s.mp4\">\u003Cb>▶ Download Full 1344x768 Video (MP4)\u003C\u002Fb>\u003C\u002Fa>\n      \u003C\u002Ftd>\n    \u003C\u002Ftr>\n    \u003Ctr>\n      \u003Ctd>\n        \u003Cul>\n          \u003Cli>\u003Cb>Sampling:\u003C\u002Fb> 4 steps (\u003Ccode>euler\u003C\u002Fcode> \u002F \u003Ccode>simple\u003C\u002Fcode>)\u003C\u002Fli>\n          \u003Cli>\u003Cb>Attention:\u003C\u002Fb> 75% video sparsity top-k (sparse DiT)\u003C\u002Fli>\n          \u003Cli>\u003Cb>Speedup:\u003C\u002Fb> \u003Cb>2.24x faster\u003C\u002Fb> than VDN; \u003Cb>9x faster\u003C\u002Fb> than native H3\u003C\u002Fli>\n          \u003Cli>\u003Cb>Peak VRAM:\u003C\u002Fb> ~13.5 GB\u003C\u002Fli>\n        \u003C\u002Ful>\n      \u003C\u002Ftd>\n      \u003Ctd>\n        \u003Cul>\n          \u003Cli>\u003Cb>Sampling:\u003C\u002Fb> 8 steps (\u003Ccode>er_sde\u003C\u002Fcode> \u002F \u003Ccode>beta\u003C\u002Fcode>)\u003C\u002Fli>\n          \u003Cli>\u003Cb>Attention:\u003C\u002Fb> Hybrid (windowed dense + linear delta state)\u003C\u002Fli>\n          \u003Cli>\u003Cb>Speedup:\u003C\u002Fb> ~3.5x faster than native H3\u003C\u002Fli>\n          \u003Cli>\u003Cb>Peak VRAM:\u003C\u002Fb> ~18.5 GB\u003C\u002Fli>\n        \u003C\u002Ful>\n      \u003C\u002Ftd>\n    \u003C\u002Ftr>\n  \u003C\u002Ftbody>\n\u003C\u002Ftable>\n\n---\n\n## 📊 Comprehensive Performance Comparison (RTX 4090)\n\n| Pipeline | Attention Mechanism | Steps \u002F Sampler | 1st Run Latency | Warm Wall Time | Peak VRAM |\n| :--- | :--- | :--- | :--- | :--- | :--- |\n| **MiniMax H3 Native Dense** | Dense Global Softmax | 50 steps \u002F Euler | ~650s | **~650s (10.8 min)** | ~22.5 GB |\n| **H3 Turbo LoRA (Dense)** | Dense Global Softmax | 4 steps \u002F Euler | ~210s | **~190s (3.1 min)** | ~21.0 GB |\n| **Video Delta Net (VDN-H3)** | Windowed + Linear Delta | 8 steps \u002F er_sde | 213.0s | **~135s (2.2 min)** | ~18.5 GB |\n| **Ref2VA VSA (This Repo)** 🚀 | **75% Video Sparsity + Dense Prefix** | **4 steps \u002F Euler** | **95.2s** | **~72 seconds** | **~13.5 GB** |\n\n---\n\n## 🧠 Technical Architecture: Why Ref2VA was Hard for VSA\n\n[FastVideo](https:\u002F\u002Fgithub.com\u002Fhao-ai-lab\u002FFastVideo) pioneered Visual Sparse Attention (VSA) for text-to-video (T2V) by clustering video tokens into 3D spatial-temporal tiles (4×4×4) and dynamically pruning 75–90% of tiles via a learned gating projection (`to_gate_compress`).\n\nHowever, **Ref2VA (Reference-to-Video)** was widely considered incompatible with VSA because:\n1. Ref2VA prepends dynamic multimodal condition segments (reference image latents, reference audio latents, and text prompt tokens) ahead of the generated video sequence.\n2. Naive 3D tiling causes tokens from different modalities to straddle the same tile, corrupting the conditioning masks and breaking identity fidelity.\n\n### The Solution: `Ref2VAVSAGatePatch`\nOur patch resolves this with an engineered two-tier attention layout:\n1. **Multimodal Dense-Exempt Prefix**: The geometry mapper isolates text, reference image, and reference audio tokens into segment-pure tiles that are **completely exempt from top-k pruning**. Reference tokens always remain 100% dense, guaranteeing strict character identity adherence.\n2. **Video-Only Sparse Attention**: VSA top-k pruning is applied *strictly* to generated-video key tiles.\n3. **Gate Transplant**: The 50 learned `to_gate_compress` projection matrices are transplanted directly onto the Ref2VA base model (`minimax_h3_ref2va_pruned_int8_convrot.safetensors`).\n\n---\n\n## 📦 Requirements\n\n- **ComfyUI** (latest version with native MiniMax-H3 support).\n- **comfy-kitchen** installed with CUDA `sol_attn` support.\n- PyTorch 2.4+ and CUDA 12.1+.\n- NVIDIA GPU with 24GB VRAM (RTX 3090, RTX 4090, A5000, L40S, etc.).\n\n### Checkpoints & Direct Download Links\nPlace the following files in your `ComfyUI\u002Fmodels\u002F` directories:\n\n| Model Type | File Name | Destination Directory | Download Link |\n| :--- | :--- | :--- | :--- |\n| **Diffusion Model** | `minimax_h3_ref2va_pruned_int8_convrot.safetensors` | `models\u002Fdiffusion_models\u002F` | [Download from Comfy-Org](https:\u002F\u002Fhuggingface.co\u002FComfy-Org\u002FMiniMax_H3_repackaged\u002Fresolve\u002Fmain\u002Fsplit_files\u002Fdiffusion_models\u002Fminimax_h3_ref2va_pruned_int8_convrot.safetensors) |\n| **VSA Gate** | `fasth3_vsa_gate.safetensors` | `models\u002Floras\u002F` | [Download from Hugging Face](https:\u002F\u002Fhuggingface.co\u002Fbarelymining\u002FComfyUI-MiniMax-H3-FastVideo\u002Fresolve\u002Fmain\u002Ffasth3_vsa_gate.safetensors) |\n| **Turbo LoRA (4-Step)** | `minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors` | `models\u002Floras\u002F` | [Download from LightX2V](https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fresolve\u002Fmain\u002Fminimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors) |\n| **Text Encoder** | `qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | `models\u002Ftext_encoders\u002F` | [Download from Comfy-Org](https:\u002F\u002Fhuggingface.co\u002FComfy-Org\u002FMiniMax_H3_repackaged\u002Fresolve\u002Fmain\u002Fsplit_files\u002Ftext_encoders\u002Fqwen3vl_32b_minimax_h3_nvfp4_awq.safetensors) |\n| **Video VAE** | `minimax_h3_video_vae_fp16.safetensors` | `models\u002Fvae\u002F` | [Download from Comfy-Org](https:\u002F\u002Fhuggingface.co\u002FComfy-Org\u002FMiniMax_H3_repackaged\u002Fresolve\u002Fmain\u002Fsplit_files\u002Fvae\u002Fminimax_h3_video_vae_fp16.safetensors) |\n| **Audio VAE** | `minimax_h3_audio_vae_fp32.safetensors` | `models\u002Fvae\u002F` | [Download from Comfy-Org](https:\u002F\u002Fhuggingface.co\u002FComfy-Org\u002FMiniMax_H3_repackaged\u002Fresolve\u002Fmain\u002Fsplit_files\u002Fvae\u002Fminimax_h3_audio_vae_fp32.safetensors) |\n\n> 💡 *Tip: The VSA Gate can also be extracted locally from any official FastH3 base checkpoint using the included `tools\u002Fextract_vsa_gate.py` script.*\n\n---\n\n## 🚀 Quick Start\n\n1. Clone or copy this repository into your ComfyUI `custom_nodes` directory:\n   ```bash\n   cd ComfyUI\u002Fcustom_nodes\n   git clone https:\u002F\u002Fgithub.com\u002FKablex\u002FComfyUI-Ref2VA-VSA\n   ```\n\n2. Copy the example character image into your ComfyUI input folder:\n   ```bash\n   cp ComfyUI-Ref2VA-VSA\u002Fassets\u002Fexample_character.jpg ComfyUI\u002Finput\u002F\n   ```\n\n3. Restart ComfyUI. The node appears under:\n   `FastH3\u002FVSA -> Ref2VA VSA Gate Transplant (EXPERIMENTAL)`\n\n4. Load the ready-to-use workflow from `workflows\u002Fref2va_vsa_4step_rtx4090.json` and click **Queue Prompt**.\n\n---\n\n## 🛠️ Node Wiring\n\n```\n[UNETLoader (Ref2VA INT8)]\n        │\n        ▼\n[LoraLoaderModelOnly (Turbo 4-Step LoRA, strength=1.0)]\n        │\n        ▼\n[Ref2VAVSAGatePatch (fasth3_vsa_gate.safetensors, sparsity=0.75)]\n        │\n        ▼\n[MiniMaxH3SigmaShift (shift_video=12, shift_audio=3)]\n        │\n        ├──► [BasicScheduler (steps=4, scheduler=simple)]\n        │\n        └──► [BasicGuider] ──► [SamplerCustomAdvanced (sampler=euler)]\n```\n\n---\n\n## 🧰 Extracting the Gate File (`tools\u002Fextract_vsa_gate.py`)\n\nIf you have a FastH3 VSA base checkpoint (`minimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot.safetensors`), you can extract `fasth3_vsa_gate.safetensors` yourself:\n\n```bash\npython tools\u002Fextract_vsa_gate.py \\\n    --input \u002Fpath\u002Fto\u002Fminimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot.safetensors \\\n    --output \u002Fpath\u002Fto\u002FComfyUI\u002Fmodels\u002Floras\u002Ffasth3_vsa_gate.safetensors\n```\n\n---\n\n## 📄 License & Acknowledgements\n\n- Licensed under the [Apache License 2.0](LICENSE).\n- Special thanks to the **FastVideo** team for the VSA-H3 attention concept and to the **ComfyUI** \u002F **comfy-kitchen** developers for the CUDA Sol-Attention kernel primitives.\n","这是一个为 ComfyUI 设计的自定义节点插件，用于加速 MiniMax H3 模型的 Reference-to-Video（Ref2VA）人物视频生成任务。其核心通过视频稀疏注意力（VSA）机制与极简采样步数（4步）显著提升推理效率，在 RTX 4090 上仅需约72秒即可生成5秒、24fps、1344×768分辨率的人物驱动视频，峰值显存占用约13.5GB。项目针对单卡消费级硬件优化，兼顾速度与资源效率，适用于需要快速迭代角色视频内容的AIGC创作、短视频原型开发及轻量级数字人演示等场景。",2,"2026-09-11 02:30:06","CREATED_QUERY"]