[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94624":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":9,"language":10,"languages":9,"totalLinesOfCode":9,"stars":11,"forks":12,"watchers":13,"openIssues":14,"contributorsCount":15,"subscribersCount":15,"size":15,"stars1d":15,"stars7d":15,"stars30d":16,"stars90d":15,"forks30d":15,"starsTrendScore":15,"compositeScore":17,"rankGlobal":9,"rankLanguage":9,"license":18,"archived":19,"fork":19,"defaultBranch":20,"hasWiki":21,"hasPages":19,"topics":22,"createdAt":9,"pushedAt":9,"updatedAt":23,"readmeContent":24,"aiSummary":25,"trendingCount":15,"starSnapshotCount":15,"syncStatus":14,"lastSyncTime":26,"discoverSource":27},94624,"Minimax-H3-Turbo","ModelTC\u002FMinimax-H3-Turbo","ModelTC","Distill Minimax-H3 into 4 steps",null,"Python",137,6,109,2,0,4,42.94,"Apache License 2.0",false,"main",true,[],"2026-08-24 04:01:22","# [Minimax-H3-Turbo](https:\u002F\u002Fgithub.com\u002FModelTC\u002FMinimax-H3-Turbo)\n\nMinimax-H3-Turbo provides MiniMax-H3 Turbo LoRA checkpoints, plus Diffusers\nbatch inference and ComfyUI workflows.\n\n## 1. Model specs\n\n\u003Ctable>\n  \u003Cthead>\n    \u003Ctr>\n      \u003Cth>Model\u003C\u002Fth>\n      \u003Cth align=\"center\">Tasks\u003C\u002Fth>\n      \u003Cth align=\"center\">Training\u003Cbr>resolution\u003C\u002Fth>\n      \u003Cth align=\"center\">Training shifts\u003Cbr>(video \u002F audio)\u003C\u002Fth>\n      \u003Cth align=\"center\">Distillation\u003Cbr>steps (NFE)\u003C\u002Fth>\n      \u003Cth align=\"center\">Recommended inference\u003Cbr>steps (NFE)\u003C\u002Fth>\n    \u003C\u002Ftr>\n  \u003C\u002Fthead>\n  \u003Ctbody>\n    \u003Ctr>\n      \u003Ctd>\n        \u003Cstrong>FL2VA Turbo 4-step v0.1\u003C\u002Fstrong>\u003Cbr>\n        \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fblob\u002Fmain\u002Fminimax_h3_fl2v_turbo_4step_v0.1.safetensors\">Diffusers\u003C\u002Fa> ·\n        \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002FKijai\u002FMiniMax-H3_comfy\u002Fblob\u002Fmain\u002Floras\u002Fminimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy.safetensors\">ComfyUI\u003C\u002Fa>\n      \u003C\u002Ftd>\n      \u003Ctd align=\"center\">FL2VA \u002F T2VA\u003C\u002Ftd>\n      \u003Ctd align=\"center\">544p\u003Cbr>\u003Csub>mixed aspect ratio\u003C\u002Fsub>\u003C\u002Ftd>\n      \u003Ctd align=\"center\">12 \u002F 3\u003C\u002Ftd>\n      \u003Ctd align=\"center\">4\u003C\u002Ftd>\n      \u003Ctd align=\"center\">4\u003C\u002Ftd>\n    \u003C\u002Ftr>\n    \u003Ctr>\n      \u003Ctd>\n        \u003Cstrong>FL2VA Turbo 8-step v1.0\u003C\u002Fstrong>\u003Cbr>\n        \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fblob\u002Fmain\u002Fminimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors\">Diffusers\u003C\u002Fa> ·\n        \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fblob\u002Fmain\u002Fminimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors\">ComfyUI\u003C\u002Fa>\n      \u003C\u002Ftd>\n      \u003Ctd align=\"center\">FL2VA \u002F T2VA\u003C\u002Ftd>\n      \u003Ctd align=\"center\">544p\u003Cbr>\u003Csub>mixed aspect ratio\u003C\u002Fsub>\u003C\u002Ftd>\n      \u003Ctd align=\"center\">12 \u002F 3\u003C\u002Ftd>\n      \u003Ctd align=\"center\">8\u003C\u002Ftd>\n      \u003Ctd align=\"center\">8 \u002F 4\u003C\u002Ftd>\n    \u003C\u002Ftr>\n    \u003Ctr>\n      \u003Ctd>\n        \u003Cstrong>FL2VA Turbo 4-step v1.0 768p\u003C\u002Fstrong>\u003Cbr>\n        \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fblob\u002Fmain\u002Fminimax_h3_fl2v_turbo_4step_v1.0_768p_bf16.safetensors\">Diffusers\u003C\u002Fa> ·\n        \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fblob\u002Fmain\u002Fminimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors\">ComfyUI\u003C\u002Fa>\n      \u003C\u002Ftd>\n      \u003Ctd align=\"center\">FL2VA \u002F T2VA\u003C\u002Ftd>\n      \u003Ctd align=\"center\">768p\u003Cbr>\u003Csub>1344x768\u003C\u002Fsub>\u003C\u002Ftd>\n      \u003Ctd align=\"center\">6 \u002F 3\u003C\u002Ftd>\n      \u003Ctd align=\"center\">4\u003C\u002Ftd>\n      \u003Ctd align=\"center\">4\u003C\u002Ftd>\n    \u003C\u002Ftr>\n    \u003Ctr>\n      \u003Ctd>\n        \u003Cstrong>Ref2VA Turbo 4-step v0.1\u003C\u002Fstrong>\u003Cbr>\n        \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fblob\u002Fmain\u002Fminimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors\">Diffusers\u003C\u002Fa> ·\n        \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fblob\u002Fmain\u002Fminimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors\">ComfyUI\u003C\u002Fa>\n      \u003C\u002Ftd>\n      \u003Ctd align=\"center\">Ref2VA\u003C\u002Ftd>\n      \u003Ctd align=\"center\">544p\u003Cbr>\u003Csub>mixed aspect ratio\u003C\u002Fsub>\u003C\u002Ftd>\n      \u003Ctd align=\"center\">12 \u002F 3\u003C\u002Ftd>\n      \u003Ctd align=\"center\">4\u003C\u002Ftd>\n      \u003Ctd align=\"center\">4\u003C\u002Ftd>\n    \u003C\u002Ftr>\n  \u003C\u002Ftbody>\n\u003C\u002Ftable>\n\n\n### Note on shift\n\nFor `NFE = N`, define the N transformer evaluation points on the unshifted grid as\n`q_i = (N - i) \u002F N`, where `i = 0, 1, ..., N - 1`.\n\nFor example, with `NFE = 4`, `video shift = 12`, and `audio shift = 3`, the shared\ngrid is `q = [1, 0.75, 0.5, 0.25]`, giving video sigma\n`[1, 0.9730, 0.9231, 0.8000] -> 0` and audio sigma\n`[1, 0.9000, 0.7500, 0.5000] -> 0`; each list therefore uses exactly four NFEs.\n\n### Note on reference-image resizing\n\nThe three reference-image resizing policies used by our workflows are based on\nthe `ref_image_size` implementation described in [ComfyUI's MiniMax H3 R2V reference-image sizing guidance](https:\u002F\u002Fdocs.comfy.org\u002Ftutorials\u002Fvideo\u002Fminimax\u002Fminimax-h3#prompting-tips-3):\n\n\u003Ctable>\n  \u003Cthead>\n    \u003Ctr>\n      \u003Cth align=\"center\">Mode\u003C\u002Fth>\n      \u003Cth>Behavior\u003C\u002Fth>\n      \u003Cth align=\"center\">Scale factor\u003Cbr>\u003Csub>before 32-pixel rounding\u003C\u002Fsub>\u003C\u002Fth>\n    \u003C\u002Ftr>\n  \u003C\u002Fthead>\n  \u003Ctbody>\n    \u003Ctr>\n      \u003Ctd align=\"center\">\u003Ccode>match\u003C\u002Fcode>\u003C\u002Ftd>\n      \u003Ctd>Matches the reference pixel area to the target canvas while preserving the reference aspect ratio. It never upscales a smaller reference.\u003C\u002Ftd>\n      \u003Ctd align=\"center\">\u003Ccode>min(1, sqrt(target_area \u002F ref_area))\u003C\u002Fcode>\u003C\u002Ftd>\n    \u003C\u002Ftr>\n    \u003Ctr>\n      \u003Ctd align=\"center\">\u003Ccode>max\u003C\u002Fcode>\u003C\u002Ftd>\n      \u003Ctd>Preserves the reference aspect ratio and only scales down references whose short edge exceeds 2048 pixels.\u003C\u002Ftd>\n      \u003Ctd align=\"center\">\u003Ccode>min(1, 2048 \u002F ref_short_edge)\u003C\u002Fcode>\u003C\u002Ftd>\n    \u003C\u002Ftr>\n    \u003Ctr>\n      \u003Ctd align=\"center\">\u003Ccode>diffusers\u003C\u002Fcode>\u003C\u002Ftd>\n      \u003Ctd>Preserves the reference aspect ratio and forces the short edge to 2048 pixels, matching the original Diffusers behavior.\u003C\u002Ftd>\n      \u003Ctd align=\"center\">\u003Ccode>2048 \u002F ref_short_edge\u003C\u002Fcode>\u003C\u002Ftd>\n    \u003C\u002Ftr>\n  \u003C\u002Ftbody>\n\u003C\u002Ftable>\n\nAll three policies keep the reference aspect ratio, use the H3 resolution grid\n(dimensions rounded to multiples of 32), and avoid cropping the reference\ncontent. In our distillation training, we use `match`, so the reference-image\npixel budget follows the target training resolution.\n\nThe Ref2VA inference entry point in this repository exposes the same three\npolicies through `--reference-resize-mode` and defaults to `match`. Passing\n`--reference-resize-mode diffusers` restores the original Diffusers behavior\n(the fixed 2048-pixel short edge). For our distilled models, **we recommend\nselecting `match`** so inference uses the same resizing policy as training.\n\n## 2. Diffusers setup and inference\n\nSee [DIFFUSERS_SETUP_AND_INFERENCE.md](DIFFUSERS_SETUP_AND_INFERENCE.md) for\nenvironment setup, checkpoint downloads, test JSON files, and single- or\nmulti-GPU inference commands.\n\n## 3. ComfyUI inference\n\nSee [COMFYUI_SETUP_AND_INFERENCE.md](COMFYUI_SETUP_AND_INFERENCE.md) for\nComfyUI requirements, model installation, inputs, prompts, and run instructions.\n\n### Example workflows\n\nReady-to-import graphs are in [example_workflows](example_workflows\u002F). The T2VA\nand I2VA graphs default to **FL2VA Turbo 8-step v1.0**; the Ref2VA graph uses\n**Ref2VA Turbo 4-step v0.1**.\n\n| Workflow | Task | Default resolution |\n|---|---|---|\n| [video_minimax_h3_t2v_lightx2v_turbo.json](example_workflows\u002Fvideo_minimax_h3_t2v_lightx2v_turbo.json) | T2VA (text-to-video + audio) | 960×544 (`16:9`, `0.5` MP) |\n| [video_minimax_h3_i2v_lightx2v_turbo.json](example_workflows\u002Fvideo_minimax_h3_i2v_lightx2v_turbo.json) | I2VA \u002F FL2VA (image-to-video + audio) | 864×480 (`16:9`, `0.4` MP) |\n| [video_minimax_h3_ref2v_lightx2v_turbo.json](example_workflows\u002Fvideo_minimax_h3_ref2v_lightx2v_turbo.json) | Ref2VA (reference-to-video + audio) | 960×544 (`16:9`, `0.5` MP) |\n\nAll graphs wrap the same MiniMax-H3 subgraph. T2VA leaves `first_frame` \u002F\n`last_frame` unconnected; I2VA connects a `LoadImage` to `first_frame`, with\n`last_frame` optional for first\u002Flast-frame interpolation. Ref2VA connects one\nor more reference images through the reference-input branch.\n\nFor detailed workflow inputs and execution steps, see\n[COMFYUI_SETUP_AND_INFERENCE.md](COMFYUI_SETUP_AND_INFERENCE.md).\n\n## 4. Roadmap\n\n1. Improve the visual quality and consistency of Ref2VA and FL2VA Turbo.\n\n## 5. Acknowledgements\n\nSome Ref2VA test cases and reference assets are adapted from public showcases on\nthe [Hailuo website](https:\u002F\u002Fhailuoai.video\u002F) and from the [MiniMax-H3 discussion\non Hugging Face](https:\u002F\u002Fhuggingface.co\u002Flightx2v\u002FMinimax-h3-Turbo\u002Fdiscussions\u002F29).\nWe thank the community contributors for sharing their test assets and prompts.\n\nSpecial thanks to the\n[MiniMax-AI\u002FMiniMax-H3](https:\u002F\u002Fgithub.com\u002FMiniMax-AI\u002FMiniMax-H3) project and the\nMiniMax team for open-sourcing the MiniMax-H3 model.\n","Minimax-H3-Turbo 是一个面向视频生成的轻量化模型优化项目，通过知识蒸馏将 MiniMax-H3 模型压缩至仅需 4 步（NFE）即可完成高质量推理。核心提供 FL2VA（帧-语言到视频）、T2VA（文本到视频）和 Ref2VA（参考图到视频）三类 Turbo LoRA 微调权重，并原生支持 Diffusers 批量推理与 ComfyUI 可视化工作流。模型训练分辨率为 544p 或 768p，兼顾效率与画质，适用于资源受限环境下的快速视频生成任务，如原型验证、AIGC 工具链集成及实时创意辅助场景。","2026-08-13 02:30:03","CREATED_QUERY"]