[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94466":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":9,"language":10,"languages":9,"totalLinesOfCode":9,"stars":11,"forks":12,"watchers":13,"openIssues":14,"contributorsCount":15,"subscribersCount":15,"size":15,"stars1d":15,"stars7d":15,"stars30d":16,"stars90d":15,"forks30d":15,"starsTrendScore":15,"compositeScore":17,"rankGlobal":9,"rankLanguage":9,"license":9,"archived":18,"fork":18,"defaultBranch":19,"hasWiki":20,"hasPages":18,"topics":21,"createdAt":9,"pushedAt":9,"updatedAt":22,"readmeContent":23,"aiSummary":24,"trendingCount":15,"starSnapshotCount":15,"syncStatus":25,"lastSyncTime":26,"discoverSource":27},94466,"ComfyUI-MiniMax-H3-Guide","ethanfel\u002FComfyUI-MiniMax-H3-Guide","ethanfel","Guided MiniMax H3 prompt preparation node for ComfyUI",null,"Python",130,17,103,1,0,22,42.97,false,"main",true,[],"2026-08-24 04:01:22","# ComfyUI MiniMax H3 Prompt Guide\n\nA dependency-free ComfyUI node pack that turns a rough video idea into the prompt structure expected by MiniMax H3. It separates endpoint frames from full-reference media, assigns explicit roles instead of guessing from file type, stores reusable image\u002Faudio Reference Sheets under ComfyUI user data, selects the H3 prompt family, and can use a loaded Qwen3-VL or Qwen3.5 CLIP to analyze visual references and enhance the result.\n\nThe node is based on MiniMax's official [base prompt guide](https:\u002F\u002Fhuggingface.co\u002FMiniMaxAI\u002FMiniMax-H3\u002Fblob\u002Fmain\u002Fdocs\u002FVIDEO_PROMPT_WRITING_GUIDE_base_en.md), [full-reference prompt guide](https:\u002F\u002Fhuggingface.co\u002FMiniMaxAI\u002FMiniMax-H3\u002Fblob\u002Fmain\u002Fdocs\u002FVIDEO_PROMPT_WRITING_GUIDE_ref_en.md), and [H3 model card](https:\u002F\u002Fhuggingface.co\u002FMiniMaxAI\u002FMiniMax-H3).\n\n## Nodes\n\n### Plan v2 ordered workflow\n\nFor new reference-heavy workflows, the Plan v2 nodes provide a typed semantic\nchain instead of asking one large form or an LLM to infer what each file means:\n\n    MiniMax H3 Project Setup (Plan v2)\n        -> optional Foley Target for video-to-audio generation\n        -> Image \u002F Video \u002F Audio Reference nodes\n        -> optional Subject Binding nodes\n        -> optional Character Replacement nodes\n        -> Shot 1\n        -> optional Attach Keyframe \u002F Attach Motion nodes for Shot 1\n        -> optional Dialogue Event nodes\n        -> Shot 2 -> its optional attachments ...\n        -> MiniMax H3 Prompt Merge (Plan v2)\n        -> optional Structured Prompt Enhancer (Plan v2)\n        -> optional Apply Structured Prose (Plan v2)\n        -> optional Prompt Review Gate (Plan v2)\n        -> MiniMax H3 Apply Reference Plan (Plan v2)\n        -> sampler\n\nEvery node consumes and returns one **MINIMAX_H3_PLAN_V2** value. Global\nReference nodes are accepted before the first Shot. A Shot uses one cut time;\nthe next cut or Project duration computes its end and the Shot also returns a\nstable `shot_handle`. Attach Keyframe and Attach Motion consume that handle and\nassign their media to exactly that Shot without a numeric `shot_scope` field.\nTheir forwarded Plan and Shot handle allow multiple attachments before the next\nShot. Dialogue Events attach to the most recently opened Shot, allowing Prompt\nMerge to assign S1, S2, and later speaker IDs from actual vocal-event order.\nEach Dialogue Event can optionally set `start_offset_seconds`: `-1` leaves its\nplacement automatic, `0` starts at the Shot opening, and a positive value starts\nthat many seconds after the Shot begins. Prompt Merge converts the offset to an\nabsolute timeline timestamp and rejects values outside the selected Shot.\nThe continuity dropdown marks an utterance as complete, split across the next\ncut, carried from the previous cut, or interrupted by the video ending. Matching\ncross-cut events receive `\u003Cscenetrans>` in both dialogue parts; a final\ninterrupted event receives `\u003Ccutoff>`. Complete dialogue keeps the user's\npunctuation exactly as entered and never moves it outside its `\u003Cd>` tag.\n\nProject Setup selects the target by native **frame count**, not an arbitrary\ndecimal duration. Its integer control advances only through the H3 `17k+5`\ngrid (`107, 124, ... 345, 362`), while the fixed `fps` field exposes the native\n24 FPS rate. The node badge and preview derive the exact playback duration as\n`frame_count \u002F fps`, including the explicit 362-frame \u002F 15.083-second endpoint.\n\nThe three media nodes require an exact relationship:\n\n- Image Reference distinguishes reusable visible content, exact first\u002Flast\n  frames, concrete keyframes, and storyboard planning. It creates a Subject\n  only when reusable visible content is explicitly selected.\n- Video Reference distinguishes source editing, continuation, visible content,\n  motion\u002Faction transfer, and camera\u002Fcut\u002Ftemporal structure.\n- Attach Keyframe to Shot registers one image as that Shot's concrete\n  composition anchor and explicitly marks it as the opening frame, an internal\n  keyframe, or the ending frame. Attach Motion to Shot transfers one video\n  clip's action to an upstream Subject only in that Shot. In accordance with\n  the H3 guide, the physical clip keeps its native Video route while its reusable\n  pose\u002Faction\u002Fmotion compiles as a separate Subject sourced from that Video.\n  Users do not write labels or scopes manually.\n- Character Replacement maps one precisely described performer in a source-edit\n  Video to one upstream identity Subject. Its appearance policy separately\n  controls identity, body, and wardrobe, while preservation switches retain the\n  source performance and surrounding scene. Prompt Merge inserts the locked\n  mapping into every selected Shot; Shot prose does not need repeated labels.\n- Audio Reference distinguishes voice, music style, beat, sound-effect\n  texture, exact dialogue\u002Flyric content, continuity, complete copy, partial\n  copy, and broad inspiration. Voice requires a target speaker; exact content\n  requires language and transcript. Its optional paired-video handle routes\n  that audio as the soundtrack belonging to the selected Video Reference.\n  Complete-copy and continuity soundtracks must cover the same source interval\n  as their paired video within one 24 FPS frame; partial\u002Flayer references may\n  deliberately cover a shorter selected interval. A single complete-copy or\n  continuity soundtrack paired to a matching native 362-frame video may share\n  its 15.083-second padded boundary; this does not raise the standalone or\n  multi-audio 15-second limit.\n\nSubject names are human aliases such as woman, truck, or wristwatch. Prompt\nMerge assigns the final Subject\u002FPicture\u002FVideo\u002FAudio numbers, validates compact\nscopes such as 3,4 or 3-4, canonicalizes native media order, and returns:\n\n1. **h3_prompt** — an immediately usable deterministic three- or six-section\n   prompt;\n2. **rewrite_request** — prose-enhancement instructions that explicitly lock\n   labels, roles, retention, speakers, dialogue, and cut times;\n3. **plan_context** — the compiled typed plan for the structured enhancer and\n   native Apply Reference Plan adapter;\n4. **problems_report** — readiness, mode, timing, inventory, exact native\n   routes, source-video intervals, replacement preservation controls, source-cut\n   and native-grid truncation warnings, and a nonblocking 350-500-word detail\n   check for reference-generation prompts, following the official guide;\n5. **h3_length** — the Project Setup native frame length.\n\nReference Sheet remains the reusable media library: connect its selected image\nor selected audio output to the matching Plan v2 reference node, where the\nworkflow-specific role is declared.\n\n### Video-to-audio Foley with a locked picture track\n\nUse **MiniMax H3 Foley Target (Plan v2)** immediately after Project Setup when\nthe source video's pictures must stay unchanged and H3 should generate only a\nnew synchronized audio track:\n\n```text\nProject Setup\n  -> Foley Target (decoded source frames, real source FPS)\n  -> optional reference assets\n  -> Shot 1 -> Shot 2 ...\n  -> Prompt Merge -> optional Structured Enhancer \u002F Review Gate\n  -> Apply Reference Plan -> sampler\n```\n\nFoley Target carries the source frames inside `h3_plan` as **target media**, not\nas an H3 Video Reference. Apply Reference Plan obtains positive conditioning\nfrom the normal H3 prompt node, discards that node's empty video stream,\nVAE-encodes the source picture track, and constructs the joint latent as:\n\n```text\nvideo latent: source video, noise mask 0  -> preserve\naudio latent: empty target audio, mask 1  -> generate\n```\n\nDo not connect the same source to Video Reference. That would add an expensive\nRef2VA video presentation without helping the audio-only latent operation. The\ncompiler deliberately creates no `\u003CVideo N>` label or native reference route\nfor the Foley source.\n\nPrompt Foley as a sound timeline, not as a request to remake the visuals:\n\n- In Project `initial_prompt`, state the overall audio goal, such as realistic\n  production Foley with no dialogue or music.\n- Make one Shot per real source cut. In each Shot, describe a visible timing\n  anchor followed by its concrete audible result: foot contact -> heel\u002Fsole\n  impact, hand closes on fabric -> cloth rustle, glass meets table -> short\n  glass-on-wood contact. Keep the order chronological.\n- Put continuous, source-grounded ambience in `overall_soundscape` in one short\n  paragraph. Keep dialogue and diegetic sound events in their exact Shots.\n- Set `non_diegetic_music` to `N\u002FA` unless an audience-only score is genuinely\n  wanted. Avoid generic “cinematic audio” wording, invented off-screen sources,\n  and visual\u002Fcamera instructions—the source picture track is latent-locked.\n\nFor example:\n\n```text\ninitial_prompt: Generate a realistic production-Foley track synchronized to the locked source video. No dialogue and no music.\n\nShot 1: A person crosses the tiled room. Each visible heel and sole contact produces a short, dry indoor footstep at the exact contact frame; clothing produces light movement rustle during each stride.\n\nShot 2: At the visible cut, the hand sets a drinking glass on a wooden table. A brief glass-on-wood contact occurs exactly when the base touches the surface, followed by a faint settling tick.\n\noverall_soundscape: Low continuous indoor room tone remains stable beneath the synchronized footsteps, cloth movement, and object contacts.\nnon_diegetic_music: N\u002FA\n```\n\nWhen Structured Prompt Enhancer uses visual analysis, the Foley source is\nsampled as timestamped visual evidence even though it is not a reference asset.\nIts appended Foley contract permits Qwen to infer ordinary physical sounds only\nfrom visible actions while forbidding new visual facts, speech, music, or unseen\nevents.\n\nThis currently requires MiniMax H3 per-token mask support from\n[ComfyUI PR #15375](https:\u002F\u002Fgithub.com\u002FComfy-Org\u002FComfyUI\u002Fpull\u002F15375) or an\nequivalent temporary compatibility patch. A temporary model-side patch can be\nconnected between the H3 model loader and the sampler, as with **MiniMax H3\nPer-Row Mask Patch**. Apply Reference Plan cannot inspect that separate MODEL\nbranch, so it constructs the correct masks without trying to reject or approve\nthe installed patch. The wiring follows\n[Ablejones's video-to-audio recipe](https:\u002F\u002Fdiscord.com\u002Fchannels\u002F1076117621407223829\u002F1532625331960152124\u002F1535135078651400223): do not present the video as a reference; preserve video\nwith mask `0`, generate audio with mask `1`, and use only positive conditioning\nfrom the prompt node. Ablejones also notes that fully masking audio retains none\nof the original soundtrack; optional audio references guide new sound but do\nnot copy it.\n\nThe Plan v2 browser extension hides irrelevant role fields, supplies upstream\nSubject pickers, validates numeric scope syntax, shows live label\u002Froute\u002Ftiming\nbadges, fixes the first Shot at 0.000 seconds, and replaces the Shot description\nbox with an editor that opens an upstream reference menu when `\u003C` is typed.\nThese conveniences do not replace Python validation and are not required in\nAPI\u002Fheadless mode.\n\nFor per-Shot composition, keep reusable identities before the timeline, then\nchain each Shot through its own media attachments:\n\n```text\nProject -> global Subject references\n        -> Shot 1 -> Attach Keyframe\n        -> Shot 2 -> Attach Keyframe\n        -> Shot 3 -> Attach Keyframe -> Attach Motion (target Subject)\n        -> Shot 4 -> Attach Keyframe\n        -> Shot 5 -> Attach Keyframe\n        -> Prompt Merge\n```\n\nThe five keyframes become five native `ref_image_N` routes and are cited only\ninside their matching Shot fields. The example marks each one as its Shot's\nopening frame; change the attachment dropdown for an internal or ending anchor.\nMotion clips become native `ref_video_N` routes, but their visible performance\nis described by a compiler-assigned action `Subject N`, not as a standalone\nwhole-video motion relationship. H3 still limits a plan to nine pictures,\nthree videos, twelve mixed reference files, and approximately fifteen\ncumulative reference-video seconds.\n\nFor a character transfer, build the setup chain in this order:\n\n```text\nProject Setup\n  -> Image Reference: Define reusable visible content \u002F Identity or appearance\n  -> Video Reference: Source video to edit\n  -> Character Replacement\n       source_video: Video Reference.reference_handle\n       replacement_subject: the Image Reference Subject alias\n       source_character_description: the woman in the red jacket\n       shot_scope: 1-3 (or all)\n  -> Shots -> Prompt Merge\n```\n\nThe source-character description identifies the performer already present in\nthe video; it is not a prompt for the replacement character. The replacement\nSubject must be upstream and have an Identity or appearance binding. Use the\npolicy dropdown to decide whether body and wardrobe remain from the video or\ncome from the reference character.\n\nFor exact sound placement, choose the relationship on the upstream Audio\nReference, then type `\u003C` in the Shot editor and select its `\u003CAudio N>` tag.\nWrite the tag inside the sentence at the moment the sound occurs, for example:\n\n```text\nEach visible impact produces the texture referenced by \u003CAudio 1>, synchronized with contact.\n```\n\n`shot_scope` validates which Shots may use the audio; it does not decide sentence\nplacement. A nonverbal audio tag placed in Shot prose suppresses the compiler's\ngeneric `overall_soundscape` fallback for that reference, so the authored Shot\nsentence remains its exact temporal location. Untagged references retain the\nglobal fallback for backward compatibility.\n\nThe optional **Structured Prompt Enhancer (Plan v2)** gives Qwen the complete\ncompiled scene in one request: the valid H3 context, all references and roles,\nevery Shot, audio metadata\u002Ftranscripts, timing, routes, and the compiler report.\nIts `enhancement_mode` defaults to **Intent-locked expansion**. In that mode,\nPython keeps every original prose field verbatim and asks Qwen only for short\nShot, compatible camera, and soundscape addenda. Attempts to replace global\nintent, visual style, music, or a previously blank camera instruction are\nignored and reported. Choose **Creative expansion** only when a complete prose\nrewrite with additional presentation choices is wanted.\nCompiler-owned dialogue lines are represented by locked placeholders instead\nof copyable H3 markup; Python restores their exact speaker, wording, voice, and\ndelivery afterward. When visual analysis is enabled, the same request also\ncontains image pixels and timestamped video samples. Audio waveforms are never\npresented as something Qwen can understand; audio meaning remains explicit\nmetadata.\n\nQwen returns only a versioned JSON object. In intent-locked mode its strings are\ntreated as addenda and composed onto the separately retained source; in creative\nmode they are complete replacements for the editable fields. The node's\n`editable_prose` output is always the final, complete validated JSON after that\ncomposition. Python then reconstructs the complete H3 prompt and verifies that\nlabels, reference roles, retention, native routes, speaker order, exact\ndialogue, Shot order, and cut times did not change. Invalid, incomplete, or\ncollapsed model output falls back to the deterministic compiler draft. The node\nexposes:\n\n1. the rebuilt `enhanced_prompt` and valid `editable_prose` JSON;\n2. the matching compiled `enhanced_plan_context`;\n3. both the editable `base_system_prompt` and actual\n   `effective_system_prompt` with its appended locked contract;\n4. the complete `llm_prompt` and a validation\u002Fmodel-residency report.\n\nThe enhancers never synchronously unload a connected complete Qwen checkpoint.\nComfyUI's normal memory manager keeps it resident and reclaims it when another\nmodel needs the VRAM. The serialized `offload_after_generation` switch remains\nfor workflow compatibility, defaults to off, and is safely ignored when enabled.\nThe temporary MiniMax generation tail still unloads itself after every request.\n\nUse **Apply Structured Prose (Plan v2)** after a text editor when the returned\nJSON should be refined manually and recompiled without running Qwen again.\nThe side-node Generation Tail Loader remains supported.\n\nUse **Prompt Review Gate (Plan v2)** for the final human check immediately\nbefore native conditioning. It is prompt-only: connect the matching prompt and\n`plan_context` from Prompt Merge, the structured enhancer, or Apply Structured\nProse. The gate pauses the queued job, opens the full H3 prompt in a large text\neditor, and offers **Approve & continue**, **Restore input**, revision history,\nand **Reject run**. Generation resumes from the same queue after approval; no\nimage, video, or audio must be selected again.\n\nThe editor permits descriptive scene, camera, ambience, and music prose changes.\nIt rejects edits to compiler-owned sections, exact reference labels, retention,\nShot markers\u002Forder, cut timestamps, dialogue tags\u002Fwords, character-replacement\ninstructions, and media routes. Approved text is bound to the connected plan\nand checked again by Apply Reference Plan. History is stored under ComfyUI user\ndata and contains prompt text and hashes only—never reference media or tensors.\n`Pass through without pausing` leaves the gate in a workflow while disabling\nthe interactive stop; the current prompt is still displayed in the editor.\n`timeout_seconds` controls the blocking limit in pause mode: `0` waits\nindefinitely, while a positive value (for example, `60`) forwards the original,\nunedited prompt automatically after that many seconds.\n\nUse **Inline Prompt Override (Plan v2)** for quick A\u002FB experiments without the\ninteractive pause. Insert it between Prompt Merge (or the Structured Enhancer \u002F\nApply Structured Prose) and Apply Reference Plan, connect both the prompt and\nmatching `plan_context`, then paste a complete experimental H3 prompt into\n`override_prompt`. Clear the field or disable `use_override` to bypass while\nkeeping the experiment in the workflow. The node applies the same descriptive-\nprose validation and plan-bound approval as the Review Gate; compiler-owned\nsections, reference labels, Shot structure\u002Ftiming, dialogue, and media routes\ncannot be overridden. It does not add the experiment to Review Gate history.\n\n```text\nWithout Qwen: Prompt Merge -> Prompt Review Gate -> Apply Reference Plan\nWith Qwen:    Prompt Merge -> Structured Enhancer -> Prompt Review Gate\n                                                        -> Apply Reference Plan\nExperiment:   Prompt Merge -> Inline Prompt Override -> Apply Reference Plan\n```\n\n**Apply Reference Plan (Plan v2)** is the native handoff. Connect `h3_prompt`\nand `plan_context` from the same Prompt Merge, Structured Prompt Enhancer, Apply\nStructured Prose, Inline Prompt Override, or Prompt Review Gate result, plus the\nofficial H3 CLIP\u002Fvideo VAE and optional audio VAE. The node verifies that the\npair still matches, automatically routes stored media as endpoint frames or\ncanonical Ref2VA dictionaries, and delegates conditioning to ComfyUI's installed\n`MiniMaxH3ImageToVideo` or `MiniMaxH3ReferenceToVideo` implementation. It\nreturns native positive conditioning and the joint AV latent. Reference audio\nrequires the audio VAE; text-only, endpoint, and reference-free Foley plans do\nnot. For Foley, it replaces the native empty video latent with the encoded\nsource and applies the per-stream masks automatically.\n\nThe adapter does not duplicate ComfyUI's encoder. It checks the installed\nnative call signature and fails with an actionable compatibility message when\nthat API changes. `adapter_report` states the selected mode, required checkpoint\nfamily, target size, length, native implementation, and every applied route.\n\n### Workflow presets and migration\n\nReady-to-open examples live in `example_workflows\u002F` and appear in ComfyUI's\nworkflow template browser. The text-only prompt builder is preconfigured for\nAPP mode; reference starters keep their semantic spine visible so files, roles,\nShots, and dialogue can be inspected before generation. The identity-and-voice\nexample also includes the current Ref2VA model, sampler, joint video\u002F\naudio decode, and Save Video path, with Apply Reference Plan replacing manual\nreference-socket wiring. It places Prompt Review Gate directly before Apply\nReference Plan to demonstrate the prompt-only pause while all media stays in\n`plan_context`.\n\n**Workflow status:** `MiniMax H3 Plan v2 - Animate Keyframe with Motion\nReference.json` is the finished reference preset. Every other example is visibly\nmarked **WIP** on its Project node and first canvas group while its task-specific\nmedia path and defaults are refined. The WIP generation examples already share\nthe finished preset's resolution, model optimization, preview, sampling, joint\ndecode, and save layout; the prompt-builder APP intentionally remains prompt-only.\n\n`MiniMax H3 Plan v2 - Video to Audio Foley.json` is the audio-generation\nstarter. It loads one video, carries its decoded frames through Foley Target,\ncompiles a sound-oriented Shot timeline, and feeds the automatically masked AV\nlatent into Apply Reference Plan. Match Apply Reference Plan's width and height\nto the source aspect ratio to avoid stretching; the node handles the exact\npixel resize and H3 frame-grid padding. The WIP includes the FL2VA loader,\nrequired per-row mask patch, sampler, joint video\u002Faudio decode, and Save Video\ntail so the locked picture track and generated Foley audio remain synchronized.\n\n`MiniMax H3 Plan v2 - Video to Audio Foley with Sound Reference.json` adds a\nclean footstep clip as a standalone `Sound-effect texture` reference. It shows\nthe important distinction between the locked target video and `\u003CAudio 1>`:\nthe audio clip supplies transient\u002Fmaterial character, while the Shot sentence\nplaces `\u003CAudio 1>` at the exact visible foot-contact event and explicitly keeps\ntiming tied to the source video. Because this is Ref2VA conditioning, the\nexample also connects the MiniMax H3 audio VAE and should be sampled with the\nRef2VA checkpoint family.\n\n`MiniMax H3 Plan v2 - Video to Audio Foley with Multiple Sound and Voice\nReferences.json` is the WIP multi-reference variant. It derives an on-screen\nSubject from the locked source frame, uses two independent sound-effect texture\nreferences for footsteps and secondary contact detail, and binds a third audio\nreference to that Subject's voice for one explicit Dialogue Event. Its PDD\u002FLoRA\nsampling path is retained from the working graph. The resized source frames are\nmuxed directly with the generated audio, so the final picture track stays\nsource-derived and no unused video decode branch is present.\n\n`MiniMax H3 Plan v2 - Character Replacement.json` is the WIP replacement-only\ngeneration preset. Set Project duration to the source video's duration,\nload the replacement identity image and a 2–15 second source video with audio,\nthen identify exactly one source performer in Character Replacement. The image\nis routed as identity-only evidence, the video supplies the complete timeline,\nand its paired soundtrack is reused as the complete synchronized output track.\nThe included Shot is a no-cut placeholder. If the source video contains cuts,\nduplicate and chain one Shot node per real source shot at the exact source cut\ntimestamps. H3's guide requires chronological shot boundaries; the compiler\nreports a warning for a one-Shot source edit because Prompt Enhancer deliberately\ncannot invent missing cuts after compilation.\nThe review gate feeds Apply Reference Plan, the official H3 video and audio\nVAEs, sampling, joint decode, and Save Video without manual reference wiring.\n\n`MiniMax H3 Plan v2 - Five Shot Keyframe Composition.json` demonstrates the\nshot-composition chain directly: five Shots, one loaded keyframe attached to\neach Shot as its opening frame, and one optional motion clip whose reusable\naction Subject is transferred to the established identity Subject only in Shot\n3. Replace the placeholder media and prose, then review the compiled plan before\nthe included Ref2VA generation, joint decode, and Save Video tail.\n\n`MiniMax H3 Plan v2 - Animate Keyframe with Motion Reference.json` is the\ncomplete one-Shot generation preset for retargeting a motion clip onto a supplied\nopening keyframe. A reusable character Picture defines stable identity, a second\nPicture anchors the Shot-opening pose and composition, and a 24-FPS Video supplies\nonly pose progression, body mechanics, cadence, and action timing. The workflow\nuses Ref2VA rather than exact endpoint conditioning so all three roles can coexist\nin one native call. Prompt Merge is set to the compact low-token style and feeds\nPrompt Review, Apply Reference Plan, sampling, joint decode, and Save Video.\n\n`MiniMax H3 Plan v2 - Video Extension with Audio Continuity.json` is a WIP\none-pass character-transfer continuation example. Load a replacement-character\nimage and one 2–10 second source video with audio. The image defines the\nreplacement Subject; the video is registered as `Source video to continue`; and\nCharacter Replacement maps one precisely described source performer to that\nSubject from the first source-derived frame through the continuation. Attaching\nCharacter Replacement to a continuation source deliberately compiles a combined\n`reference generation + video editing + video continuation` target: it first\nrecreates the source timeline with that performer replaced, then continues the\nedited endpoint. The Picture supplies identity and appearance only and is\nexplicitly forbidden from becoming an opening frame, standalone shot, or\nanimated segment. The loader's audio is registered as\n`Audio continuity` and paired to that exact Video Reference. The compiled prompt\nkeeps audio synchronized with the source-derived portion, then requires it to\ndevelop forward without restarting, replaying, repeating, or looping after the\nendpoint. Project duration is the total edited-plus-continued output duration,\nso it must be longer than the loaded source video. This is one H3 generation,\nnot a second replacement pass; the saved result already contains both the\ncharacter-transferred source-derived portion and its continuation.\nAs with replacement-only editing, represent every real cut in the source-derived\nportion with its own chained Shot node and exact cut timestamp; the single Shot\nin the template is only a no-cut placeholder.\n\nFor existing graphs, see [Migrating existing workflows to Plan v2](MIGRATION_TO_PLAN_V2.md).\nOld node IDs remain registered so saved workflows load, but the monolithic\nPrompt Guide, Target Timing\u002FShot chain, visual\u002Faudio context builders, and\nfree-form enhancer now include **Legacy** in their library names. Reference\nSheet and Generation Tail Loader remain supported components.\n\n### Legacy: MiniMax H3 Prompt Guide\n\n`MiniMax H3 Prompt Guide` appears under `MiniMax H3\u002FPrompting`. It produces:\n\n1. `h3_prompt` — a deterministic, structured pre-LLM draft. When a final Visual Reference `reference_context` is connected, the Guide derives role-correct Subject grouping, direct Picture\u002FVideo rows, retention relationships, and the H3 family from that context. Run the draft through the enhancer when the creative notes are short or visual analysis would help.\n2. `rewrite_request` — a self-contained instruction for an LLM or Context-IR-style rewrite step. This is recommended when the starting notes are short because the official guide expects a detailed chronological description.\n3. `mode_report` — the selected mode and checkpoint, the reason for the selection, Ref2VA limits, and warnings about contradictory options.\n4. `h3_length` — the requested duration rounded upward to native ComfyUI's `17k+5` frame grid at 24 FPS. In a simple workflow it can feed the official H3 node directly. When a final Visual Reference context returns to the Guide, use the upstream Target Timing node described below; the Guide then echoes the same resolved value without creating a graph cycle.\n\nNo model, API key, or extra Python dependency is required for the guide itself.\n\n### Legacy: MiniMax H3 Target Timing\n\nUse **MiniMax H3 Target Timing** whenever a final Visual Reference\n`reference_context` will feed the Prompt Guide. It is especially important for\nvideo references because it resolves duration before video preparation and\nexposes:\n\n1. `timing_context` — connect to `Prompt Guide.timing_context`; it carries the\n   requested\u002Feffective duration and any connected Shot chain.\n2. `h3_length` — connect to every video Visual Reference `h3_length` input and\n   to the official H3 node's `length` input.\n3. `timing_report` — the selected timing source and native `17k+5` result.\n\nThis keeps every edge pointing downstream: Target Timing prepares the length,\nvideo references use it to trim their analysis\u002Fnative batches, and only then\ndoes the final reference context reach the Guide. Do not feed\n`Prompt Guide.h3_length` back into a video Visual Reference that contributes to\nthe Guide's own `reference_context`.\n\n### Legacy: MiniMax H3 Prompt Enhancer (Qwen LLM)\n\nThis node follows ComfyUI's native **Generate Text** execution model. Connect\n`h3_prompt` and optionally `mode_report` from the guide node. Its `CLIP` input\naccepts a complete generation-capable Qwen3-VL or Qwen3.5 model, or MiniMax H3's\nnormal 50-layer conditioning CLIP plus the optional 50–63 generation tail:\n\n```text\nMiniMax H3 Target Timing\n    timing_context ──────────────────────> Prompt Guide.timing_context\n    h3_length ─────┬─────────────────────> each video Visual Reference.h3_length\n                  └─────────────────────> native H3.length\n\nfinal Visual Reference.reference_context ─┬─> Prompt Guide.reference_context\n                                          └─> Prompt Enhancer.reference_context\nfinal Reference Sheet Audio.audio_context ──> Prompt Guide.audio_context\n\nPrompt Guide.h3_prompt ────┐\nPrompt Guide.mode_report ──┼─> MiniMax H3 Prompt Enhancer ─> enhanced_prompt\nstandard CLIPLoader.CLIP ──┤                              ├─> system_prompt\nlegacy optional IMAGE ─────┤                              ├─> llm_prompt\nGeneration Tail Loader ────┘                              └─> enhancer_report\n```\n\nIt produces:\n\n1. `enhanced_prompt` — Qwen's cleaned candidate H3 prompt. Generation residue is\n   removed, but the text is not silently rewritten after decoding; review\n   `enhancer_report` for structural warnings before generation.\n2. `system_prompt` — the resolved base enhancer instructions, exposed so they can be reused, inspected, or edited.\n3. `llm_prompt` — the serialized text\u002Fchat portion sent to Qwen. Pixel tensors and MiniMax reference blocks are tokenizer inputs and therefore are not embedded in this string. An external multimodal LLM must receive the pixels separately in its own visual-token format.\n4. `enhancer_report` — the resolved H3 family, generation status,\n   compatibility\u002Ffallback details, and structural H3 warnings.\n\nThe full base system prompt is visible in the node's editable `system_prompt` widget. If that widget is blank, the built-in default is restored. Exact unmodified defaults serialized by older releases are upgraded automatically; any customized prompt is preserved. Sampling controls match the important controls from ComfyUI's Generate Text node: maximum generated tokens, deterministic or sampled decoding, temperature, top-k, top-p, min-p, repetition\u002Fpresence penalties, seed, and thinking mode.\n\nThe optional `image` remains as a compatibility route for one context image.\nFor labeled pictures, multiple images, or video understanding, use the Visual\nReference chain below. Visual context helps Qwen write the prompt; it never\nsilently turns a picture into an endpoint frame or replaces the media inputs\non the native H3 node.\n\nFor a complete generative Qwen3-VL or Qwen3.5 CLIP, leave the enhancer's\noptional `clip_tail` socket disconnected. Qwen3.5 uses ComfyUI's normal\n`CLIPLoader` and the native Generate Text contract; its 4B model is a practical\ngeneral enhancer choice. The native MiniMax 32B text encoder is\nthe deliberately truncated conditioning model described below; loading that\ncheckpoint does not by itself create a complete 32B LLM.\n\nFor MiniMax H3's bundled conditioning CLIP, add **MiniMax H3 Generation Tail\nLoader**, select\n`qwen3vl_32b_minimax_h3_generation_tail_50_63_int8_convrot.safetensors` in\nthe loader, and connect its `clip_tail` output to the enhancer. The loader\npasses a lightweight descriptor and consumes no VRAM by itself. During\nenhancement, the enhancer reuses the connected embedding, vision tower, and\nlanguage layers 0–49, loads only layers 50–63 plus the final norm and LM head,\nthen unloads that tail when generation finishes. Tail KV caches and embeddings\nare released before the managed unload and CUDA cache flush. The connected\n50-layer CLIP is never merged or modified and remains suitable for official H3\nconditioning. The connected CLIP remains under ComfyUI's normal memory manager;\nthe enhancer does not synchronously unload it after decoding. The legacy\n`offload_after_generation` widget is retained only so older workflows continue\nto load, and enabling it no longer forces a connected-model unload.\nIf the truncated CLIP is connected without the side loader, enhancement is\nsafely skipped and the manual prompt is returned unchanged.\n\nThe tail loader accepts only the published split layout. Its chunked LM head\nsupports ComfyUI tensor-wise INT8 scalar\u002Fper-row scales and rejects other\nquantized layouts explicitly. The complete model does not have to fit in VRAM:\nthe base and tail are both registered with ComfyUI's managed patchers, so\nDynamicVRAM streams\u002Fcaches weights on demand and legacy Normal VRAM can\npartially load them. This means a card below 32 GB may run the enhancer when it\nhas enough VRAM for the largest active layer, KV cache, vision tensors, and\nruntime headroom, plus enough system RAM for offloaded weights; that full path\nhas not yet been hardware-verified below 32 GB. It will be much slower because\nautoregressive generation revisits every language layer for each token.\n`--highvram` and especially `--gpu-only` defeat this low-VRAM behavior; with\n`--gpu-only`, the configured load and offload devices are identical.\n\nHigh `nvidia-smi` usage on a larger card does not itself mean full residency is\nrequired: ComfyUI opportunistically uses available VRAM and may retain allocator\ncache. Explicit post-generation cleanup returns only the transient tail; the\nconnected CLIP remains under ComfyUI's model manager. A real INT8 tail artifact\nhas now been smoke-tested for successful text generation; available hardware\nstill determines practical speed and maximum visual\u002Fprompt context.\n\nDownload the INT8 tail from\n[`ethanfel\u002FQwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8-ConvRot`](https:\u002F\u002Fhuggingface.co\u002Fethanfel\u002FQwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8-ConvRot)\nand place it under `ComfyUI\u002Fmodels\u002Ftext_encoders\u002FMiniMax-H3\u002F`.\n\nConnect `enhanced_prompt` to ComfyUI's official **MiniMax H3 Image to Video** or **MiniMax H3 Reference to Video** node. Those nodes encode the prompt and attach the correct AV latent plus any keyframe\u002Freference VAE latents and media metadata. The enhancer deliberately does not emit a separate `CONDITIONING` output because it would duplicate the official node for T2VA and be incomplete for image\u002Freference tasks.\n\n### Persistent Reference Sheets\n\nA **Reference Sheet** is a reusable media library entry rather than a\ncharacter-only profile. One sheet can describe a person, object, outfit,\nlocation, style, voice, sound, or a mixed project collection. Saved sheets live\noutside the custom-node repository under:\n\n```text\nComfyUI\u002Fuser\u002Fdefault\u002Fminimax_h3\u002Freference_sheets\u002F\n└── \u003Csheet-name>--\u003Cshort-id>\u002F\n    ├── manifest.json\n    ├── images\u002F\n    └── audio\u002F\n```\n\nSet `MINIMAX_H3_REFERENCE_SHEETS_DIR` before starting ComfyUI only when a\ndifferent library root is required. Each manifest is versioned and contains a\nUUID, descriptions, tags, relative media paths, and SHA-256 checksums. Connected\nComfyUI `IMAGE` tensors are saved as PNG and connected `AUDIO` tensors are saved\nas WAV, making the sheet independent of its creation workflow and original\ninput filenames. Create never overwrites another sheet; Update requires\n`confirm_update` and merges the selected sheet atomically. Every existing media\nasset is preserved, newly connected non-duplicate media is appended with a new\nstable key, and byte-identical connections are skipped.\n\nBuild, display, and save a sheet with one integrated node:\n\n```text\nLoad Image.IMAGE ───> Reference Sheet.image_1\nLoad Image.IMAGE ───> Reference Sheet.image_2      optional\nLoad Audio.AUDIO ───> Reference Sheet.audio_1      optional\n\nReference Sheet\n    operation: Create new\n    sheet_name: reusable display name\n    └─ embedded thumbnail\u002Faudio gallery\n```\n\nConnect up to four image sources and three audio sources directly. An image\nbatch is expanded into separate saved pictures, up to H3's nine-image limit.\nQueue once to create the sheet. Later choose `Load existing`, select the sheet,\nand click the desired thumbnail or audio player in the embedded gallery. The\nselection is carried in `reference_sheet`; users never type or remember an\ninput filename or internal asset key. For audio, `audio_start_seconds` chooses\nthe offset and the `audio_duration_seconds` numeric field chooses an exact 2–15 second output\nwindow. Trimming is non-destructive: the player and saved WAV remain complete,\nwhile `selected_audio` and the legacy sheet-audio output carry only the selected\nsegment. Duplicate the sheet node when different segments of one saved clip are\nneeded in the same workflow.\n\nUpdate with no connected media changes\nmetadata only and preserves the saved media. Update with connected media appends\nto the saved collection after `confirm_update` is enabled; it never treats the\ncurrently connected inputs as a complete replacement list.\n\nUse saved assets directly with Plan v2:\n\n```text\nReference Sheet.selected_image -> Plan v2 Image Reference.image\nReference Sheet.selected_audio -> Plan v2 Audio Reference.audio\n```\n\nThe Plan v2 reference node assigns the actual workflow relationship, retention,\nscope, speaker\u002Flayer binding, and native route. Duplicate Reference Sheet when\nseveral saved assets must be selected independently. The old Reference Sheet\nVisual\u002FAudio Reference context nodes remain registered for existing Prompt\nGuide workflows and are labeled Legacy.\n\nThe structured Qwen enhancer can analyze the selected image after it enters the\nPlan v2 image inventory. It does not analyze the audio waveform: audio meaning\ncomes from the exact Audio Reference metadata, while native H3 receives the real\n`AUDIO` value through Apply Reference Plan. Reference Sheet stores images and\nstandalone audio; use Video Reference for decoded video frame batches.\n\n### Legacy: visual references, roles, and native routing\n\nUse one **MiniMax H3 Enhancer Visual Reference** node per picture or reference\nvideo. `previous_context` records assets in chain order. The backend then\nnumbers pictures and videos independently and recommends sockets in native H3\ncategory order: pictures first, then videos, while preserving chain order\nwithin each category. A separate **MiniMax H3 Visual Reference Role** chain\nassigns one or more semantic jobs to a single media file:\n\n```text\nRole: identity ─> Role: clothing ─> Visual Reference.role_bindings\n                                       │ media: Picture\nPrevious Visual Reference.context ─────┤\n                                       ├─ reference_context ─> next Visual Reference.previous_context\n                                       └─ h3_media ─────────> socket recommended by routing_report\n\nfinal Visual Reference.reference_context ─┬─> Prompt Guide.reference_context\n                                          └─> Prompt Enhancer.reference_context\n```\n\nNew Visual Reference nodes start with `Unassigned - choose a reference role`.\nBefore running, either select one simple role in the compatibility\n`reference_role` dropdown or connect a completed `role_bindings` chain. Use\nrole nodes when one asset has several roles, when several assets should provide\nevidence for one Subject, or when retention\u002Fshot mapping must be explicit. Fan\nthe final Visual Reference `reference_context` out to both the Prompt Guide and\nPrompt Enhancer. The Guide deterministically writes the role-correct Subject or\ndirect Picture\u002FVideo rows; Qwen then analyzes the supplied pixels and expands\nthe creative description without being asked to invent the role mapping.\n\nThe role fields mean:\n\n| Field | Meaning |\n| --- | --- |\n| `reference_role` | What content the asset provides: endpoint, identity, object, scene, style, keyframe, storyboard, motion, temporal structure, edit source, or continuation source. The `Unassigned - choose a reference role` new-node default must be replaced before execution. |\n| `retention` | One official visible marker: `fully_preserved`, `partially_preserved`, `attribute_transfer`, or `weak_reference`. Auto uses full preservation for identity\u002Fobject\u002Fscene, weak reference for style\u002Fstoryboard\u002Ftemporal structure, and attribute transfer for action\u002Fmotion. |\n| `content_group` | A stable user key for reusable visible content. Give bindings on different files the same key when the Guide should combine them as evidence for one `\u003CSubject N>`. |\n| `transfer_target` | Required whenever retention resolves to `attribute_transfer`, including Auto action\u002Fmotion bindings; names a different upstream Subject that receives the attribute or motion. |\n| `shot_scope` | Optional Shot numbers: `3`, `3,4`, `3-4`, or `all`. Older wording such as `Shot 3` remains supported. Leave blank when the location is not known instead of inventing Shot 1. |\n| `notes` | What to preserve, transfer, ignore, or change for this binding. |\n\nTwo route families are intentionally exclusive:\n\n- **Endpoint context:** `Exact first frame` and\u002For `Exact last frame` stays in\n  I2VA\u002FL2VA\u002FFL2VA. The report maps each `h3_media` output to native **MiniMax\n  H3 Image to Video** `first_frame` \u002F `last_frame`. Two endpoint pictures are\n  analyzed in native Picture 1\u002FPicture 2 order.\n- **Ref2VA context:** reusable content, concrete keyframes, storyboards, motion,\n  temporal structure, edit sources, and continuation sources receive recommended\n  native **MiniMax H3 Reference to Video** `ref_image_N` \u002F `ref_video_N` inputs.\n\nThe backend `routing_report` is the authoritative description of the intended\nmapping. It cannot create native-node links: connect every `h3_media` output to\nthe listed socket yourself. Canvas output labels are convenience hints,\nespecially when role chains, reroutes, or bypassed nodes are present. Do not\nmix endpoint and Ref2VA roles in one context chain.\n\nThe media paths deliberately have different representations:\n\n- **Picture passed to H3:** the original picture is unchanged. In Ref2VA,\n  native `ref_image_size=match|max` remains authoritative.\n- **Video passed to H3:** the source batch is resampled to 24 FPS, optionally\n  truncated to connected `h3_length`, and rounded downward to native H3's\n  `17k+5` reference grid. Set `source_fps` to the real batch rate and connect\n  Target Timing's `h3_length` to every video reference node when the final\n  context also feeds the Guide. Use the Guide's output only in a legacy path\n  where doing so cannot form a cycle.\n- **Generic-Qwen analysis:** a reduced long-edge copy with configurable\n  `analysis_fps` and frame cap, sampled only from the effective native clip.\n- **MiniMax-Qwen analysis:** a separate fixed-2-FPS sequence, matching native\n  MiniMax temporal pairs. It is not affected by the generic frame cap.\n\nReference videos must be 2–15 seconds before native alignment and total at most\n15 seconds. The report shows both source and effective duration, and Qwen's\nvisual evidence excludes the discarded tail. Free-form text can still mention\ndiscarded events, so review the candidate prompt when the source was trimmed. A\n48-frame 24-FPS source, for example, becomes 39 native frames and four MiniMax\nsamples at 0.0, 0.5, 1.0, and 1.5 seconds.\n\nH3's `\u003CSubject N>` is reusable visible content, not a synonym for a person. It\nmay represent an object, environment, style, action, expression, or pose. From\nthe connected context, the Guide cites a picture\u002Fvideo used only as reusable\nSubject evidence inside that Subject's definition without adding an unnecessary\nstandalone definition\u002Fretention row. Concrete frames remain `\u003CPicture N>`;\nedit\u002Fcontinuation\u002Fwhole-video temporal sources remain `\u003CVideo N>`. The enhancer\ncan expand those definitions from visual evidence and checks the resulting\nstructure, but the explicit bindings remain authoritative. Review\n`enhancer_report` and the candidate text before generation. Reference Sheet\naudio uses a separate typed context because the enhancer model receives its\nsaved text description rather than the waveform.\n\n### Planning multiple shots\n\nAdd one **MiniMax H3 Shot** node per shot and connect each `shot_plan` output to\nthe next node's `previous_shots` input. With a chained Visual Reference context,\nconnect the final Shot to Target Timing so the complete plan stays upstream:\n\n```text\nMiniMax H3 Shot (0.000–2.500)\n    shot_plan ─> MiniMax H3 Shot (2.500–4.250)\n                    shot_plan ─> MiniMax H3 Shot (4.250–6.000)\n                                    shot_plan ─> Target Timing.shot_plan\n\nTarget Timing.timing_context ─> Prompt Guide.timing_context\nTarget Timing.h3_length ──────┬─> every video Visual Reference.h3_length\n                              └─> native H3.length\n```\n\nWithout a connected `reference_context`, the older direct route remains valid:\nconnect the final Shot to `Prompt Guide.shot_plan`, and use\n`Prompt Guide.h3_length` downstream.\n\nEach node has float `start_time` and `end_time` controls with millisecond\nsteps, a shot description, per-shot camera direction, and transition. The\nchain rejects gaps, overlaps, reversed ranges, a first shot that does not start\nat zero, and times above 15 seconds. Its final `end_time` is the requested\nduration. Target Timing—or the Guide in the legacy direct path—rounds that\nduration to native H3 frames and extends the last described shot through the\neffective playback end. Camera instructions are written as natural shot prose,\nnever as a `Camera direction:` metadata label.\n\nThe Prompt Guide's `shot_and_timing_plan` text widget remains available under\nadvanced controls for old workflows or a quick manual plan. It now parses the\nsame common syntax into real H3 markers, for example:\n\n```text\nShot 1, 00:00-00:02.500: medium entrance.\nShot 2, cut at 00:02.500: close-up reaction.\nShot 3, cut at 00:04.250: wide ending.\n```\n\nNumbers, gaps, ranges, descriptions, and cut order are validated. A connected\nShot chain takes priority. Final-frame alignment always cites a Shot marker\nthat actually exists.\n\n## Choosing the right route\n\n| What you want | Mode | Checkpoint |\n| --- | --- | --- |\n| Generate from text | T2VA | H3-Base-FL2VA |\n| Animate an exact first frame | I2VA | H3-Base-FL2VA |\n| Land on an exact final frame | L2VA | H3-Base-FL2VA |\n| Connect exact first and last frames | FL2VA | H3-Base-FL2VA |\n| Use appearance, style, motion, video editing\u002Fcontinuation, or audio references | Ref2VA | H3-Base-Ref2VA |\n\nThe important distinction is the role of the asset, not merely its file type:\n\n- A picture used as the exact first or last frame is an endpoint anchor.\n- A picture used only for a character's appearance or scene style is a reference-generation asset.\n- A video being modified is `video editing`; a video that only supplies motion, camera movement, cuts, or rhythm is `reference generation`.\n- Copying an audio signal is `audio reuse`; borrowing its timbre, beat, music style, or sound texture is `audio reference`.\n\nFor the common “transfer motion to an image” case, select:\n\n- `Transfer motion to a different subject`\n- `Target subject for motion transfer`\n- `Transfer its motion or action`\n\nThis creates a Ref2VA prompt where the target image keeps its visible identity and the action reference receives the fixed `attribute_transfer` relationship.\n\nWith per-asset role nodes, the equivalent explicit mapping is:\n\n| Media | Role | Content group | Retention | Transfer target |\n| --- | --- | --- | --- | --- |\n| Picture 1 | Identity or appearance | `hero` | `fully_preserved` | — |\n| Video 1 | Motion or action | `reference-motion` | `attribute_transfer` | `hero` |\n\nTo combine rather than transfer evidence, give bindings the same group. For\nexample, Picture 1 `Identity or appearance` and Video 1 `Motion or action` can\nboth use `hero`; the Guide then defines one Subject whose appearance comes from\nthe picture and whose motion evidence comes from the video, without inventing\nan unrelated second Subject. The enhancer may add details found in the media,\nbut it receives the same fixed grouping.\n\n## Reference inventory\n\nEnter one asset per line. Labels describe the native sockets to which you intend\nto connect the corresponding media:\n\n```text\nPicture 1: a red ceramic robot, front three-quarter view\nVideo 1: a dancer performing a quick clockwise spin\nAudio 1: a dry studio recording of a calm female voice\n```\n\nAngle brackets are optional. `Picture 1: ...` and `\u003CPicture 1>: ...` are\nequivalent. Labels must be positive and unique. Active gaps are rejected;\nout-of-order entries are reported in `mode_report` so you can match native\ncategory order. Unlabelled lines are retained as additional reference notes.\n\nInventory text describes expected files; it does not decide their role. A\nlisted Picture with image role `No image`, for example, stays unused and\nproduces a warning instead of silently becoming an appearance reference or\nSubject. In the legacy dropdown path, selecting a role may make the text-only\nGuide synthesize a required placeholder label to keep the draft structurally\ncomplete, including Picture 2 when a partially listed first-and-last-frame task\nneeds it. A placeholder is not a media file: `mode_report` calls it out, and you\nmust verify that every generated label has a real downstream connection.\n\nThe Guide's global image\u002Fvideo role dropdowns remain a legacy shortcut when no\n`reference_context` is connected. In that path they select one role per media\ntype and may create structurally required placeholder labels, exactly as older\nworkflows expect.\n\nFor chained references, the final Visual Reference `reference_context` is the\nsingle authoritative visual model. Connect it to both the Guide and Enhancer.\nThe Guide ignores its legacy image\u002Fvideo role dropdowns, derives Subject\ngrouping, direct Picture\u002FVideo rows, retention, task prefix, and H3 family from\nthe explicit bindings, and uses matching inventory lines only as descriptions.\nThe enhancer analyzes and expands that already aligned draft instead of\nreconciling two conflicting role models. Legacy audio still uses the Guide's\ndropdown and inventory. When a final Reference Sheet `audio_context` is\nconnected, its saved descriptions and per-workflow audio relationship replace\nthat legacy audio path and provide exact standalone native routes.\n\n## Example: edit a video and keep its soundtrack\n\nSet `how_video_is_used` to `Directly edit the source video` and `how_audio_is_used` to `Reuse the complete audio signal`. The node selects Ref2VA and starts the summary with:\n\n```text\n[video editing + audio reuse] The target video is an edited version of \u003CVideo 1>.\n```\n\nThe output also distinguishes the video retention marker from the audio\n`fully_copy` marker. The Guide states that no new layer may be added and warns\nwhen its combined dialogue\u002Ftext field could imply a new vocal signal. The\nenhancer instructs Qwen to keep the copy exclusive, and its structural check\nrequires applicable audio sections to cite the copied label and state that\nexclusivity. Free-form target descriptions and custom system prompts still\ncannot be semantically proven compatible; review them and use partial copy when\nthe target adds or replaces sound.\n\n## Install\n\nClone or copy this folder into `ComfyUI\u002Fcustom_nodes\u002F` and restart ComfyUI:\n\n```bash\ncd ComfyUI\u002Fcustom_nodes\ngit clone https:\u002F\u002Fgithub.com\u002Fethanfel\u002FComfyUI-MiniMax-H3-Guide\n```\n\nFor new work, add nodes from **MiniMax H3 → Plan v2**, beginning with Project\nSetup and ending with Prompt Merge plus Apply Reference Plan. Reference Sheet\nappears under **MiniMax H3 → Reference Sheets**, and Generation Tail Loader\nremains under **MiniMax H3 → Prompting**. Other Prompting\u002Fcontext nodes include\nLegacy in their displayed names for old-workflow compatibility. The previous\nfilename-based Reference Sheet Image Asset, Audio Asset, and Library builder\nnodes remain disabled.\n\nThis release expects a ComfyUI build containing native MiniMax H3 support\n(introduced by ComfyUI commit `57500fc5bc92`). Update ComfyUI if the official\n**MiniMax H3 Image to Video** \u002F **MiniMax H3 Reference to Video** nodes or the\nMiniMax tokenizer are absent.\n\n## Practical notes\n\n- H3's requested output range is 4–15 seconds. Native ComfyUI rounds upward to\n  `17k+5` frames, so the effective value shown in the prompt\u002Freport can be\n  slightly longer (for example, 7.25 seconds becomes 175 frames \u002F 7.292 seconds).\n- H3 Ref2VA policy allows up to 9 images, 3 videos, 3 audio clips, and 12 media\n  files in total. Plan v2 validates the complete mixed inventory before\n  conditioning.\n- H3 policy requires each reference video\u002Faudio clip to be 2–15 seconds and\n  limits each media type to 15 seconds total. The only native-grid exception is\n  one complete-copy or continuity soundtrack paired to a matching 362-frame\n  video; both may cover the padded 15.083-second source interval. Video\n  Reference and Audio Reference validate each asset when registered, and Prompt\n  Merge validates their totals. Audio Reference errors include the cumulative\n  duration and each active clip's duration; use the Reference Sheet numeric trim\n  fields to fit several selected segments below the shared limit.\n- H3 policy does not allow reference audio as the sole media input. Prompt Merge\n  rejects that plan before Apply Reference Plan can run. Foley is the explicit\n  exception supported by the native conditioning path: its locked target video\n  supplies the visual stream while optional audio-only Ref2VA guidance supplies\n  sound characteristics.\n- Native Ref2VA ordering is pictures first; then each enabled video soundtrack\n  `\u003CAudio N>` immediately before its `\u003CVideo N>`; then standalone audio. Audio\n  and video labels are independently numbered, so equal numbers do not imply a\n  pairing. Audio Reference creates a paired `ref_video_audio_N` route only when\n  its optional Video Reference handle is explicitly connected.\n- Native Image to Video stretches a first frame to the target canvas and\n  center-cover-crops a last frame. Match endpoint aspect ratio to output\n  width\u002Fheight when exact composition is important.\n- Apply Reference Plan wires Plan v2 media automatically. Expert users may keep\n  the official native conditioning node and connect sockets manually according\n  to Prompt Merge's route report.\n- A Foley target intentionally replaces the complete audio stream. It rejects\n  `Copy complete signal` and `Copy selected part or layers`; choose a timbre,\n  sound-texture, beat, continuity, or broad reference when sonic guidance is\n  needed. Preserving or partially regenerating an original soundtrack requires\n  an explicit audio-time mask workflow outside this full-Foley helper.\n- Community testing reports that stochastic\u002FSDE sampling can damage H3 audio.\n  If Foley is noisy or muted, first test a deterministic\u002FODE path (or `eta=0`)\n  before changing the prompt.\n- For an external\u002Fgeneral LLM node, use the structured editable-prose contract\n  and Apply Structured Prose. A raw full-prompt rewrite can no longer be paired\n  safely with the native adapter unless it is represented by matching compiled\n  plan data.\n- The included enhancer can use the same MiniMax H3 CLIP as conditioning when the Generation Tail Loader is connected. Managed partial loading is designed to support sub-32-GB cards, but that hardware tier remains unverified and the 32B autoregressive pass can be substantially slower when weights stream from system RAM. Leave DynamicVRAM enabled and avoid `--highvram` \u002F `--gpu-only` for that use case.\n- See [AUDIT_REPORT.md](AUDIT_REPORT.md) for the source-grounded findings,\n  compatibility decisions, and remaining limitations.\n\n## Test\n\n```bash\npytest -q\n```\n","这是一个为ComfyUI设计的MiniMax H3视频生成模型专用提示词结构化工具包，无需外部依赖。它将粗略的视频创意转化为H3模型严格要求的结构化提示格式，支持显式标注图像\u002F视频\u002F音频参考素材的角色（而非依赖文件类型自动推断），内置参考素材管理、多镜头时序编排、关键帧与运镜绑定、对话事件时间对齐及结构化提示增强等功能，并可选集成Qwen3-VL\u002FQwen3.5 CLIP进行视觉参考分析。适用于需要高精度控制MiniMax H3视频生成流程的专业创作者和AI视频工作流开发者。",2,"2026-08-09 02:30:05","CREATED_QUERY"]