Guides
How to Kill CUDA OOM on Wan 2.2 and LTX-2 Without Buying a 5090
· RenderBob team
Out-of-memory errors on video workflows are the single most common ComfyUI complaint of 2026, and most of them are solvable without new hardware. Work this list in order before you reach for a credit card.

Out-of-memory errors on video workflows are the single most common ComfyUI complaint of 2026, and most of them are solvable without new hardware. Work this list in order before you reach for a credit card.
1. Disable pinned memory
A wave of OOM and host-buffer crashes on Wan 2.2 and MiniMax H3 after recent updates trace back to pinned memory, which locks host RAM for faster transfers but creates system-RAM pressure. Launching with --disable-pinned-memory resolves many of these with no measurable speed loss. This is the first thing to try.
2. Lower the two variables that matter most: resolution and frame count
VRAM scales hardest with these, and the classic LTX-2 crash appears specifically at 1080p above ~200 frames when the workflow enters its upscale pass. Generate at a lower base resolution and upscale as a separate step, or split a long clip and chain it.
3. Turn on adaptive/dynamic VRAM and offloading
Recent ComfyUI builds added adaptive model loading to reduce OOMs and Windows shared-memory spilling. Combine with --lowvram (keeps the text encoder on CPU), --async-offload, and disabling live previews (--preview-method none) to reclaim working memory.
4. Use quantized weights
NVFP4 on RTX 50-series or FP8 elsewhere can cut VRAM roughly 40–60%. A workflow that OOMs in full precision often runs comfortably quantized.
5. Watch system RAM, not just VRAM
Several 2026 regressions were RAM problems in disguise: models not offloading, pagefile thrashing, and SSD wear. If your GPU shows headroom but the job still dies, look at RAM and your pagefile before blaming the card.
6. Save latents and split the graph
Rendering the sampling pass, saving latents, then running the upscale pass separately sidesteps the VRAM spike where both stages briefly coexist. It is inelegant and it works.
If you have worked this list and a job still will not fit (a genuinely heavy hero shot at full resolution and length), you have hit a hardware ceiling, not a configuration one. Route that job to a cloud node with more VRAM. Do not re-spec the whole farm for it. Pay for extra capacity on the jobs that need it, and keep everything else on hardware you already own.
More from the blog
- A Risk Ladder for AI in Documentary
Screenweaver's 8 October guide ranks five documentary uses of AI by risk, from archive restoration to a synthetic face, and pairs them with EU disclosure rules now in force.
- The B-Roll Gap: Generated, Selected, or Shot
Every cut eventually needs a shot that does not exist. AI can generate it or search a library for it, and stock libraries are answering with different rules. Here is how to choose.