MiniMax H3 can run locally in ComfyUI, including on a single RTX 3090, but the practical limit is not just GPU VRAM. Model files, the text encoder, system RAM, storage, offloading behavior, resolution, and clip length all affect whether a workflow loads and how long it takes.
For the least-friction path, update ComfyUI to version 0.30.0 or later, start with its native MiniMax H3 template, use the optimized Comfy-Org weights, and test a short low-resolution clip before attempting the native 768-pixel-short-edge canvas. A documented community RTX 3090 setup produced a 5-second 832×480 clip in about 4 minutes 26 seconds and a 15-second clip in about 23 minutes 17 seconds. Those are third-party results from one Linux configuration, not SeeAPI benchmarks or universal speed promises.
If your goal is to create instead of tune a local stack, you can open MiniMax H3 on SeeAPI and use text, frames, or multimodal references without downloading the weights.
MiniMax H3 ComfyUI requirements at a glance
Hardware tier | Practical expectation | Recommended starting point |
|---|---|---|
24 GB VRAM, such as RTX 3090 | The best-documented single-GPU local tier. Optimized weights and offloading are still important. | T2V or I2V, 832×480, about 5 seconds, then scale gradually. |
12–16 GB VRAM | Possible with quantized weights and aggressive offloading, but system RAM and storage pressure rise sharply. | Use the official template, shortest clip, low megapixel target, and close other GPU workloads. |
8 GB or less | A technical experiment rather than a comfortable default. ComfyUI reports that optimized variants can run on consumer GPUs through dynamic offloading. | Expect long loads and generation times; consider cloud or SeeAPI if iteration speed matters. |
48 GB+ VRAM | More headroom for larger canvases and fewer offload stalls, but still not a guarantee of fast long clips. | Validate the workflow first, then increase resolution, duration, or reference count one variable at a time. |
The original MiniMax H3 weights are much larger than a 24 GB card. Consumer-GPU workflows depend on pruning, quantization, and dynamic offloading; a successful launch does not mean the entire model fits in VRAM.
What is actually open for local use?
MiniMax describes H3 as an omni-modal generation system that understands text, images, video, and audio and produces video with native stereo audio. Its public specifications cover 4–15 second output, 24 fps, a 768-pixel default short edge, and several landscape, square, and vertical aspect ratios.
However, the complete official system has three parts:
H3-Context-IR interprets and restructures complex multimodal instructions. MiniMax says this hosted orchestration system is not part of the open-source release.
H3-Base is the locally deployable generation model. Separate FL2VA and Ref2VA checkpoints handle text/keyframe generation and multimodal reference generation.
H3-Regenerate-2K uses the original context and the 768p result to regenerate more detailed 2K output. MiniMax says this module is not yet open-sourced.
That distinction matters. A local H3-Base result and the full hosted 2K workflow are not identical products. For a predictable first local test, treat 768p-class output as the baseline rather than assuming every advertised 2K capability is available entirely offline.
Step 1: update ComfyUI and choose the native workflow
MiniMax H3 support is built into ComfyUI 0.30.0 and later. After updating, open Template Library → Video and select one of these workflows:
MiniMax H3 T2V for prompt-only video generation.
MiniMax H3 I2V for one image or first-and-last-frame control.
MiniMax H3 R2V for identity, style, motion, camera, voice, or sound references.
Use the native template before installing third-party node packs. It gives you known node wiring and the filenames expected by the current ComfyUI implementation. If nodes such as MiniMaxH3ImageToVideo, MiniMaxH3ReferenceToVideo, or the MiniMax tokenizer are missing, the first fix is usually to update ComfyUI rather than rebuild the graph manually.
Step 2: download the correct model files
For text-to-video and image-to-video, the official ComfyUI template expects these files:
Component | Filename | Folder |
|---|---|---|
Diffusion model |
|
|
Text encoder |
|
|
Video VAE |
|
|
Audio VAE |
|
|
Optional turbo LoRA |
|
|
Reference-to-video uses the Ref2VA diffusion model and its matching turbo LoRA instead of the FL2VA pair. Do not mix the two task families and expect the graph to load correctly.
The four core T2V/I2V files occupy roughly 41 GB in one current optimized distribution before the optional LoRA. Leave additional free disk space for downloads, cache files, outputs, and future model revisions.
Step 3: start with a conservative RTX 3090 configuration
A community-tested RTX 3090 configuration used Ubuntu, a single 24 GB card, approximately 32 GB of system RAM, additional swap, ComfyUI 0.30.1, and the optimized INT8/NVFP4 files. The launch command included:
python3 main.py \
--listen 0.0.0.0 \
--port 8188 \
--disable-pinned-memory \
--fp16-intermediatesIn that setup, disabling pinned memory allowed the operating system to page memory instead of trapping too much RAM in non-pageable buffers. --fp16-intermediates reduced intermediate tensor size. These flags are a practical troubleshooting starting point, not a requirement for every OS or ComfyUI build.
For your first queue:
Select the T2V template.
Use 832×480 or the template's fast preview preset.
Choose a clip around 5 seconds.
Keep the default sampler and step count.
Generate once before enabling turbo mode, Sage Attention, extra references, or a larger canvas.
After that baseline succeeds, change only one variable at a time. Increasing duration and resolution together makes it much harder to identify whether a failure comes from VRAM, system RAM, the attention sequence, or decode.
Step 4: understand resolution and duration
ComfyUI's current guide defines H3's native 16:9 canvas as 1344×768, or about 0.98 megapixels. Resolutions should stay on a multiple-of-32 grid. The duration control also snaps to H3's 17k + 5 frame grid at 24 fps.
Longer clips cost disproportionately more compute. In one RTX 3090 report, moving from 124 frames to 362 frames increased sampling time far more than the 2.9× frame-count increase. That behavior is consistent with attention becoming more expensive as the sequence grows, but it should not be treated as a universal scaling formula.
If 15 seconds fails, do not immediately assume the installation is broken. Return to a short preview, confirm the output has both moving frames and audio, then raise duration in steps.
Step 5: write a prompt that includes sound
H3 models video and stereo audio together, so the prompt should describe the scene, camera, dialogue, sound effects, and music as one coordinated brief.
Scene: A compact robotics workshop at night. A small repair drone wakes on a workbench.
[Shot 1] Slow push-in as indicator lights turn on and the drone raises its head.
[Shot 2] The camera pans right while the drone rolls toward a half-built mechanical bird.
Dialogue:
Technician, off camera: "Let's try one more time."
Overall soundscape: quiet ventilation, small servo movements, one soft electrical click.
Non-diegetic music: restrained analog synth pulse, low volume.Keep the first test structurally simple. Once it works, the MiniMax H3 prompt guide explains how to improve shot timing, camera direction, reference roles, and native-audio cues.
Common MiniMax H3 ComfyUI problems
Symptom | Likely cause | What to try first |
|---|---|---|
H3 nodes or template are missing | Old ComfyUI build | Update to 0.30.0 or later and restart. |
Model dropdown is empty | Wrong folder, filename, or an inaccessible symlink | Check the exact model directories and container mounts. |
Process is killed while loading | System RAM or swap exhaustion | Close other workloads, use the NVFP4 text encoder, disable pinned memory where appropriate, and watch host memory. |
CUDA out-of-memory during sampling | Canvas, duration, or another GPU process is too large | Reduce megapixels and frames; free GPU memory; test T2V before R2V. |
Output is black, frozen, or silent | Decode, workflow, or file issue | Inspect frames and the audio stream instead of trusting that an MP4 file exists. |
R2V produces weak identity or motion | Reference roles are ambiguous | State which image controls identity and which video controls motion; keep the first test small. |
Turbo mode will not load | Missing or mismatched LoRA | Install the FL2VA or Ref2VA LoRA that matches the chosen workflow. |
When local ComfyUI is the right choice
Run MiniMax H3 locally when you need graph-level control, repeatable automation, offline asset handling, custom nodes, or the ability to inspect every stage. A 24 GB GPU is a useful enthusiast tier, but local ownership also means managing roughly 40 GB of core files, RAM pressure, CUDA and PyTorch compatibility, long queues, and model updates.
Use an online workflow when you need to validate a prompt quickly, compare 768P and 2K output, avoid hardware tuning, or let a team create from the same interface. SeeAPI currently exposes MiniMax H3 text-to-video, image-to-video, and reference-to-video workflows with 4–15 second output and native stereo audio.
Final recommendation
For an RTX 3090, start with the native ComfyUI T2V template, optimized FL2VA weights, NVFP4 text encoder, 832×480, and a short clip. Confirm motion and audio before increasing duration, resolution, or reference count. Treat system RAM and storage as part of the model's memory budget, and separate the open H3-Base workflow from the hosted 2K regeneration pipeline.
If that setup cost is larger than the creative task, generate with MiniMax H3 on SeeAPI and return to local ComfyUI when you need deeper graph control.







