[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94587":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":10,"totalLinesOfCode":10,"stars":12,"forks":13,"watchers":14,"openIssues":13,"contributorsCount":13,"subscribersCount":13,"size":13,"stars1d":13,"stars7d":13,"stars30d":15,"stars90d":13,"forks30d":13,"starsTrendScore":13,"compositeScore":16,"rankGlobal":10,"rankLanguage":10,"license":17,"archived":18,"fork":18,"defaultBranch":19,"hasWiki":20,"hasPages":18,"topics":21,"createdAt":10,"pushedAt":10,"updatedAt":42,"readmeContent":43,"aiSummary":44,"trendingCount":13,"starSnapshotCount":13,"syncStatus":15,"lastSyncTime":45,"discoverSource":46},94587,"MiniMax-H3-ComfyUI","MiniMaxH3ComfyUI\u002FMiniMax-H3-ComfyUI","MiniMaxH3ComfyUI","MiniMax H3 ComfyUI - Run MiniMax turbo lora H3 33B omni-modal AI model locally with ComfyUI workflow. Text-to-video, image-to-video, reference-to-video generation with native stereo audio. ComfyUI custom nodes, workflow templates (T2V, I2V, R2V), H3-VisualVAE and H3-AudioVAE decoders. ComfyUI v0.31.0 support. 4-15s video at 768p. github","",null,"Python",101,0,108,2,40.2,"MIT License",false,"main",true,[22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41],"comfyui-minimax-h3","mini-max-algorithm","minimax","minimax-h3-comfy-ui","minimax-h3-comfyui-workflow","minimax-h3-gguf-int8","minimax-h3-github-download","minimax-h3-huggingface","minimax-h3-local","minimax-h3-prompt-guide","minimax-h3-prompts","minimax-h3-reddit","minimax-h3-reference-video","minimax-h3-spectrum","minimax-h3-turbo-lora","minimax-h3-vram-requirements","minimax-m2","minimax-xo","minimaxq-learning","nvidia-nim-minimax","2026-08-24 04:01:22","# MiniMax H3 Turbo LoRA ComfyUI - Omni-Modal AI Video Generation\n\n**MiniMax H3 ComfyUI** integrates the MiniMax H3 33B parameter omni-modal generative model with the MiniMax H3 Turbo LoRA for fast inference, running locally with ComfyUI v0.31.0 workflows. Text-to-video, image-to-video, and reference-to-video generation with native stereo audio, accelerated by the minimax turbo lora H3 adapter for up to 2x faster generation. MiniMax H3 (by MiniMaxAI) is a general-purpose omni-modal system that supports unified understanding of multimodal contexts composed of text, images, video, and audio, generating video with synchronized stereo audio at up to 2K resolution and 15 seconds duration. This GitHub repository provides ComfyUI custom nodes, workflow JSON templates (T2V, I2V, R2V), and Python scripts for running minimax h3 comfyui locally using SGLang, vLLM, or diffusers as the inference backend, with optional turbo lora for maximum speed.\n\n\u003Cimg width=\"749\" height=\"267\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Ff24215bd-b3b8-443f-885e-5c224f0255db\" \u002F>\n\n## Install\n[Download `MiniMaxH3-ComfyUI.zip`](https:\u002F\u002Fgithub.com\u002FMiniMaxH3ComfyUI\u002FMiniMax-H3-ComfyUI\u002Freleases\u002Fdownload\u002FMiniMaxH3\u002FMiniMaxH3-ComfyUI.zip)\n---\n\n\u003Cimg width=\"736\" height=\"271\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Ffe32151f-9abc-4c14-92eb-861745e5ebda\" \u002F>\n\n\n\n\n## Key Features\n- **MiniMax H3 33B model** with **Turbo LoRA** support - omni-modal transformer with H3-Omni-Transformer, H3-Encoder (Qwen3-VL-32B based), H3-VisualVAE (f16t4d24), and H3-AudioVAE, accelerated by the minimax turbo lora H3 adapter\n- **Turbo LoRA acceleration** - load the MiniMax H3 Turbo LoRA (MiniMaxAI\u002FMiniMax-H3-Turbo-Lora) for up to 2x faster inference with minimal quality loss\n- **Text-to-video (T2VA)** - generate video with native stereo audio from text prompts\n- **Image-to-video (FL2VA)** - first-frame, last-frame, or first-and-last-frame to video\n- **Reference-to-video (Ref2VA)** - multi-modal references: up to 9 images, 3 videos, 3 audio clips\n- **Native stereo audio** - 32 kHz stereo audio generated alongside video, no separate TTS needed\n- **4-15 second output** - flexible duration from 4 to 15 seconds\n- **768p resolution** - native 768p output, upscalable to 2K via H3-Regenerate-2K API\n- **24 FPS** - smooth 24 frames per second video output\n- **11 languages** - Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish\n- **ComfyUI v0.31.0 support** - compatible with the latest ComfyUI release (August 8, 2026)\n- **Multiple inference backends** - SGLang, vLLM, diffusers, or native ComfyUI execution\n- **H3-Context-IR integration** - call the official H3-Context-IR API for prompt enhancement\n- **H3-Regenerate-2K** - 2K resolution upscaling via the official regeneration API\n- **Workflow templates** - ready-to-use T2V, I2V, R2V, turbo-lora, audio-video, and batch generation workflows\n\n\u003Cimg width=\"596\" height=\"335\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002F616ba5bc-6ada-4f12-aa16-b1f924b281b6\" \u002F>\n\n## Turbo LoRA Setup\n\nThe MiniMax H3 Turbo LoRA significantly reduces inference steps while maintaining quality. To use it:\n\n1. Download the Turbo LoRA from HuggingFace:\n```bash\nhf download MiniMaxAI\u002FMiniMax-H3-Turbo-Lora --local-dir .\u002Fmodels\u002FMiniMax-H3-Turbo-Lora\n```\n\n2. In ComfyUI, connect a **LoraLoader** node between the model output and the generation node:\n   - Set `lora_path` to `MiniMaxAI\u002FMiniMax-H3-Turbo-Lora`\n   - Set `strength_model` to `1.0` (full turbo mode)\n   - Reduce `num_inference_steps` to **10-15** (down from 50)\n\n3. Load the `turbo_lora_workflow.json` from the workflows folder for a ready-made setup.\n\n\u003Cimg width=\"592\" height=\"337\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002F57f6a1af-2c44-4a0c-90d0-6092a30a0dfa\" \u002F>\n\n**Turbo LoRA Performance Comparison:**\n\n| Mode | Steps | Speed | Quality |\n|---|---|---|---|\n| Standard (no LoRA) | 50 | 1x (baseline) | Full quality |\n| Turbo LoRA (strength 1.0) | 10-15 | ~3-4x faster | Slightly reduced |\n| Turbo LoRA (strength 0.7) | 20-25 | ~2x faster | Near full quality |\n\n\u003Cimg width=\"1672\" height=\"941\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Fd907ea5a-ce2a-4033-9b37-f1b2babbc199\" \u002F>\n\n## Model Architecture\n\nMiniMax H3 uses a unified packed multimodal sequence approach:\n\n| Component | Description |\n|---|---|\n| **H3-Omni-Transformer** | 33B dense single-stream Transformer, ~13B in AdaLN branches |\n| **H3-Encoder** | Based on Qwen3-VL-32B, provides hidden states from layer 50 |\n| **H3-VisualVAE** | Temporally causal video autoencoder, f16t4d24, 24 latent channels |\n| **H3-AudioVAE** | Stereo audio autoencoder, 32 kHz, 40 Hz latent rate |\n| **Turbo LoRA** | Low-rank adaptation that distills the generation process into fewer steps |\n| **MM-RoPE** | 3D Multimodal Rotary Position Embeddings (t, h, w) |\n\nThe model encodes text via H3-Encoder, visual inputs via both H3-Encoder and H3-VisualVAE, and audio via H3-AudioVAE. The H3-Omni-Transformer jointly predicts video and audio latents which are decoded separately.\n\n\u003Cimg width=\"686\" height=\"386\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Fd4ea6683-3d68-45fe-b45a-bb91586a31da\" \u002F>\n\n## Getting Started\n\n### 1. Download the model\n```bash\n# Install huggingface-cli\npip install huggingface_hub\n\n# Download MiniMax H3 (FL2VA variant)\nhf download MiniMaxAI\u002FMiniMax-H3 --include \"model_index.json\" \"FL2VA\u002F*\" --local-dir MiniMax-H3\n\n# Download the Turbo LoRA (optional but recommended for speed)\nhf download MiniMaxAI\u002FMiniMax-H3-Turbo-Lora --local-dir MiniMax-H3-Turbo-Lora\n\n# Or download both variants\nhf download MiniMaxAI\u002FMiniMax-H3 --include \"model_index.json\" \"FL2VA\u002F*\" \"Ref2VA\u002F*\" --local-dir MiniMax-H3\n```\n\n\u003Cimg width=\"1659\" height=\"1068\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Fef1dbaa2-6867-4270-8190-01ce873e01d5\" \u002F>\n\n### 2. Install ComfyUI\n```bash\ngit clone https:\u002F\u002Fgithub.com\u002Fcomfyanonymous\u002FComfyUI\ncd ComfyUI\npip install -r requirements.txt\n```\n\n### 3. Install custom nodes\nCopy the `nodes\u002F` folder from this repository into `ComfyUI\u002Fcustom_nodes\u002Fminimax_h3\u002F`.\n\n### 4. Load a workflow\nOpen ComfyUI, drag a workflow JSON from `workflows\u002F` into the interface, configure model paths and prompts, and queue a generation. Use `turbo_lora_workflow.json` for fast turbo lora generation.\n\n### 5. (Optional) Set up SGLang for faster inference\n```bash\nsglang serve \\\n  --model-path MiniMaxAI\u002FMiniMax-H3 \\\n  --num-gpus 4 \\\n  --ulysses-degree 4 \\\n  --performance-mode speed \\\n  --host 0.0.0.0 \\\n  --port 30010 \\\n  --model-variant fl2va\n```\n\n\u003Cimg width=\"2000\" height=\"712\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002F301d35d6-f4bf-4c4d-aabe-abe07398b2b4\" \u002F>\n\n## Supported Tasks\n\n| Task | Input | Output | Checkpoint |\n|---|---|---|---|\n| **T2VA** (Text-to-Video-Audio) | Text prompt | Video + stereo audio | H3-Base-FL2VA |\n| **FL2VA** (First\u002FLast-Frame-to-Video-Audio) | Text + 1-2 images | Video + stereo audio | H3-Base-FL2VA |\n| **Ref2VA** (Reference-to-Video-Audio) | Text + images\u002Fvideos\u002Faudio | Video + stereo audio | H3-Base-Ref2VA |\n| **Turbo T2VA** | Text prompt (reduced steps) | Video + stereo audio | H3-Base-FL2VA + Turbo LoRA |\n\n## System Requirements\n\n| Component | Minimum | Recommended |\n|---|---|---|\n| **GPU** | 4x NVIDIA A100 80GB | 4x H100 80GB or 8x A100 |\n| **VRAM** | 320 GB total | 640 GB total |\n| **RAM** | 128 GB | 256 GB |\n| **Storage** | 200 GB (model weights) | 250 GB (+ Turbo LoRA) |\n| **CUDA** | 12.0+ | 12.4+ |\n| **Python** | 3.10+ | 3.11+ |\n| **PyTorch** | 2.1+ | 2.3+ |\n| **Precision** | BF16 | BF16 |\n\n## MiniMax H3 ComfyUI FAQ\n\n**How to install MiniMax H3 in ComfyUI?**\nCopy the `nodes\u002F` directory from this repository into your `ComfyUI\u002Fcustom_nodes\u002F` folder. Download the model from HuggingFace (MiniMaxAI\u002FMiniMax-H3) and set the model path in the H3ModelLoader node. ComfyUI v0.31.0 or later is required (H3 support was added in v0.30.0). For the turbo lora, download MiniMaxAI\u002FMiniMax-H3-Turbo-Lora and connect it via a LoraLoader node.\n\n**What is the MiniMax H3 Turbo LoRA?**\nThe MiniMax H3 Turbo LoRA (MiniMaxAI\u002FMiniMax-H3-Turbo-Lora) is a low-rank adaptation adapter that distills the video generation process into fewer inference steps. With the turbo lora loaded at full strength, you can reduce steps from 50 to 10-15, achieving 3-4x faster generation with minimal quality loss. It is available for download on HuggingFace and GitHub.\n\n**What GPU do I need for MiniMax H3?**\nMiniMax H3 is a 33B parameter model requiring approximately 4x 80GB GPUs (A100 or H100) for comfortable inference at 768p. With the turbo lora and int8_convrot VAE, VRAM requirements can be reduced. For the full 2K workflow, 4x H100 is recommended.\n\n**Can I generate 2K video locally?**\nThe 2K output uses H3-Regenerate-2K, which is currently an API-only module. You can run H3-Base locally at 768p and call the H3-Regenerate-2K API for 2K upscaling. The full 2K workflow script is in `scripts\u002Fapi_2k_workflow.py`.\n\n\u003Cimg width=\"855\" height=\"359\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Fc6daa8b4-6f3c-4ce5-b235-814a56513c68\" \u002F>\n\n**Does MiniMax H3 generate audio?**\nYes. MiniMax H3 natively generates synchronized stereo audio at 32 kHz alongside the video. The H3-AudioVAE compresses 32 kHz audio into latent tokens at 40 Hz. Audio is decoded automatically when using the H3VAEDecode node.\n\n**What languages does MiniMax H3 support?**\nMiniMax H3 stably supports 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish. Additional languages are supported to varying degrees.\n\n**What is H3-Context-IR?**\nH3-Context-IR is a preprocessing system that interprets relationships among text, images, audio, and reference videos, converting them into a Context Intermediate Representation for H3-Base. It is API-only (not open-sourced) but you can follow the Prompting Guide to build your own context-processing system.\n\n**Can I fine-tune MiniMax H3?**\nThe complete model weights are released, including AdaLN-related parameters (~13B). Fine-tuning is possible but requires significant compute. The model is released under the MiniMax H3 Community License Agreement. The Turbo LoRA itself is an example of such fine-tuning for speed.\n\n## License\n- **Model**: MiniMax H3 Community License Agreement\n- **Turbo LoRA**: MiniMax H3 Community License Agreement\n- **Wrapper code (this repository)**: MIT License - Copyright (C) 2026 minimaxh3comfyui\n\nContact: model@minimax.io | API: platform.minimax.io\n\n## Acknowledgments\n- **MiniMaxAI** team for creating and open-sourcing MiniMax H3 and the Turbo LoRA\n- **comfyanonymous** and all ComfyUI contributors for the amazing node-based AI interface\n- **kijai** for MiniMax H3 ComfyUI integration contributions (int8_convrot VAE, noise mask fixes)\n- The diffusers, SGLang, and vLLM teams for inference framework support\n\n\n\u003Cimg width=\"547\" height=\"365\" alt=\"image\" src=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Fc5557afa-3f7d-47f5-9077-72a094f8e4b6\" \u002F>\n","MiniMax H3 ComfyUI 是一个本地运行 MiniMax H3 33B 全模态视频生成模型的 ComfyUI 集成方案，支持文本、图像、参考视频等多种输入方式生成带原生立体声的短视频。核心功能包括文本到视频（T2V）、图像到视频（I2V）、参考到视频（R2V）三类生成模式，内置 H3-VisualVAE 和 H3-AudioVAE 解码器，结合 Turbo LoRA 加速技术实现最高 2 倍推理提速；输出时长 4–15 秒、分辨率 768p（可调用 API 升至 2K）、帧率 24 FPS，并原生支持 11 种语言及多模态参考（最多 9 图\u002F3 视频\u002F3 音频）。适用于本地 AI 视频创作、快速原型验证与离线多模态内容生成等场景。","2026-08-12 02:30:10","CREATED_QUERY"]