Local text-to-video is real now, but only in a narrow band of hardware. This guide uses Wan 2.2 TI2V-5B, the hybrid text/image-to-video model that ComfyUI recommends for consumer cards.
1. The VRAM reality
Read this before downloading 20 GB.
| Hardware | What actually happens |
|---|---|
| 8 GB+ VRAM | Wan 2.2 5B works — ComfyUI’s docs say it “should fit well on 8GB vram with the ComfyUI native offloading” |
| 12–16 GB | Same model, comfortably, with room for longer clips |
| 24 GB (4090) | Wan’s own README: a 5-second 720P clip “in under 9 minutes on a single consumer-grade GPU” |
| Under 8 GB | Nothing worth your evening. Use a hosted service. |
The 14B Wan models are not a consumer option — the project documents them as needing on the order of 80 GB for single-GPU inference at standard settings. Offloading spills weights into system RAM, so RAM matters here almost as much as VRAM, and a slower card means minutes-per-second-of-video, not seconds.
2. Update ComfyUI
Video nodes move fast; a template that fails to load almost always means an old build. Use Manager → Update All in the ComfyUI interface, or, for a manual install:
git pull
pip install -r requirements.txt
Brings ComfyUI and its dependencies to the current version. The desktop app updates itself.
3. Download three files
From the Comfy-Org/Wan_2.2_ComfyUI_Repackaged repository on Hugging Face, into these exact folders:
ComfyUI/models/diffusion_models/wan2.2_ti2v_5B_fp16.safetensors
ComfyUI/models/vae/wan2.2_vae.safetensors
ComfyUI/models/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors
Three separate folders — this is the step people get wrong. Wan is not a single checkpoint, so models/checkpoints stays empty here.
4. Open the template
In ComfyUI, open the Template Library and search for Wan2.2 5B. That loads the official workflow with the nodes already wired: the diffusion model loader, the UMT5 text encoder, the Wan VAE, a sampler, and a video output node.
Before running, click each loader node and confirm it points at the file you downloaded.
5. Render
Type your prompt into the positive prompt node, keep the template’s resolution and frame count for the first run, then press Ctrl + Enter.
The finished clip lands in ComfyUI/output. Watch the terminal, not the browser: the first run also loads ~20 GB from disk, so the initial render is much slower than the second.
Short prompts describing one continuous shot work best — camera move, subject, setting. Ask for a cut or a scene change and you get mush.
Next: set up ComfyUI properly first or upscale the frames afterwards.