How much VRAM does ComfyUI need? It depends on the job

ComfyUI itself needs almost no VRAM: the model you load decides. In our pipeline, 24 GB runs the face workflow and the Telegram bot’s endpoint, 48 GB makes datasets smoothly, and the course trains a character LoRA on a 96 GB RTX PRO 6000. Z-Image Turbo’s model card says it fits 16 GB cards. Our RunPod guide has the current price of each card.

On this page
  1. Why does the model, not ComfyUI, decide how much VRAM you need?
  2. How much VRAM does each task need?
  3. Is 24 GB of VRAM enough?
  4. How much VRAM does AI video need?
  5. Is more VRAM always better?
  6. How do you check how much VRAM a workflow uses?
  7. What can go wrong?

Why does the model, not ComfyUI, decide how much VRAM you need?

ComfyUI is the interface that wires models together. The VRAM goes to what it loads: the image model, the text encoder, the VAE, any LoRA, plus the image or video being worked on. A bigger model, a bigger image, a bigger batch or more video frames all need more. ComfyUI moves models out to normal RAM when VRAM runs short.

Take our face workflow. It loads three model files plus a LoRA: Z-Image Turbo (a 6-billion-parameter model, per its model card), the qwen_3_4b text encoder and the ae VAE. Those three are the files ComfyUI’s own Z-Image Turbo example uses.

When VRAM runs short, ComfyUI often offloads parts of the model to normal RAM and slows down instead of failing. On a rented GPU, slow is money. When it can’t offload enough, you get an out-of-memory error: here are the fixes.

How much VRAM does each task need?

These are the cards the course’s templates run on, as of October 2026, plus the official minimum where the model’s makers publish one. “What we use” is what’s tested; “less works?” tells you what happens below it, or what the official docs say.

Task Model What we use Less works?
Make her face Z-Image Turbo + a LoRA, 1024×1536 24 GB (RTX 4090), about 5 s per image warm Model card: fits 16 GB cards
Telegram bot (serverless) Image model + her LoRA 24 GB GPU Not tested below 24 GB
Dataset Maker, free engine Qwen-Image-based edit model on the pod 48 GB or more is smooth 32 GB fine; 24 GB offloads and is slow
Dataset Maker, paid engines Nano Banana Pro or Seedream through RunPod’s API Cheapest GPU, e.g. RTX A4000 (16 GB) Yes: the GPU barely works
Train her LoRA AI Toolkit, Krea 2 96 GB (RTX PRO 6000), 2 to 3 hours 32 to 48 GB cards need quantization + Low VRAM: a worse LoRA
Make video Wan 2.2 32 GB (RTX 5090) ComfyUI docs: the 5B model fits 8 GB with offloading
Train a video LoRA Wan 2.2 96 GB (RTX PRO 6000) Not tested on less

Two rows are easy to misread. The dataset step needs far more VRAM than making her face, because the free engine is a big edit model, not a fast image model. And with a paid engine the GPU is almost idle, so a cheap 16 GB card is the right pick.

Is 24 GB of VRAM enough?

For images, yes. Our face workflow and the Telegram bot’s serverless endpoint both run on 24 GB. For the Dataset Maker’s free engine it works but offloads, so it’s slow. For LoRA training, no: the course trains on a 96 GB RTX PRO 6000. For video, it depends on the model files.

The owner’s rule of thumb for datasets: 48 GB and up is smooth, 32 GB is fine, 24 GB is slow. Training is the one job where the course doesn’t step down: on a 32 to 48 GB card the trainer needs quantization and Low VRAM, and the LoRA comes out worse. Most people who rent should learn on a 24 GB card and move up only for the jobs that need it. Services that don’t let you choose the card, like Colab, make this harder: RunPod vs Google Colab.

How much VRAM does AI video need?

Wan’s scripts in its official repo list 24 GB for the 5B model and 80 GB for the 14B models on one GPU. ComfyUI needs far less: its Wan 2.2 docs say the 5B fits 8 GB with native offloading, and ship fp8 14B files (checked 2 Oct 2026). We generate on a 32 GB RTX 5090.

In the course, a Wan 2.2 video LoRA trains on a 96 GB RTX PRO 6000, and the finished video LoRA is about 300 MB. If you’re starting with video, see how to make AI influencer videos for which tool fits which budget before you rent the big card.

Is more VRAM always better?

No. Enough VRAM matters; extra VRAM you don’t use just costs more. A job that fits on a 32 GB card can finish cheaper there than on a 96 GB card. In the owner’s estimates for a 50-photo dataset, the RTX 5090 was cheapest per dataset (about $1.90) against the RTX PRO 6000 (about $2.70). LoRA training is the exception.

Those are estimates, not measured runs, as of October 2026. The point holds anyway: VRAM decides whether a job fits, speed and price decide what it costs. Once the job fits without offloading, pick on price per finished job, not on the biggest number.

So rent the smallest card where the job fits without offloading, and move up only when it fails; our out-of-memory guide has the upgrade path. The exception is her LoRA: there the course pays for the RTX PRO 6000, because a 32 or 48 GB card means quantization and Low VRAM, and a worse LoRA you’ll use in every photo.

How do you check how much VRAM a workflow uses?

Run nvidia-smi while the workflow is working. On a RunPod pod, open JupyterLab on port 8888, start a Terminal and type it. It shows how much memory the GPU is using against its total. If you’re near the total and the run is slow, the workflow is offloading and a bigger card will help.

Run it twice: once when the models have just loaded, once in the middle of sampling. The second number is the one that matters. To keep watching, nvidia-smi -l 2 refreshes it every two seconds; press Ctrl+C to stop.

What can go wrong?

  • You rent by the biggest model you’ve read about. Our face workflow runs on 24 GB. Rent for the job in front of you.
  • “It works” on a small card but takes forever. It’s offloading. You pay for every slow minute, so the bigger card can cost less overall.
  • You trained her LoRA on a 32 to 48 GB card. The trainer switched to quantization and Low VRAM, and the LoRA is worse. Train the one you keep on an RTX PRO 6000 (96 GB).
  • You pay for a big GPU that sits idle. With the Dataset Maker’s paid engines, the work happens in the API. A cheap card is enough.
  • The card you want is unavailable. Pick another card with the same VRAM, not a smaller one.
  • A video workflow runs out of memory. Use the smaller or fp8 model files the ComfyUI workflow is built for, or a bigger card. For Wan 2.2, ComfyUI’s Wan 2.2 page lists the fp8 files.

Questions people ask

Is 24 GB of VRAM enough for ComfyUI?
For making images, yes. Our face workflow and our Telegram bot's endpoint both run on 24 GB cards. For the Dataset Maker's free engine it works but offloads and is slow. For LoRA training, no: the course trains on a 96 GB RTX PRO 6000.
Is 16 GB of VRAM enough?
For Z-Image Turbo image generation, its model card says it fits 16 GB consumer cards. For datasets on the free engine or LoRA training with our templates, no. If you rent, a 24 GB card costs little more and gives you room.
How much VRAM do I need to train a LoRA?
The course trains on an RTX PRO 6000 with 96 GB, with Low VRAM off and no quantization. It takes 2 to 3 hours. On a smaller card like the RTX 5090, the trainer needs quantization and Low VRAM, and the LoRA comes out worse.
How do I check how much VRAM a workflow uses?
Open a terminal in JupyterLab on the pod and run nvidia-smi while the workflow runs. It shows memory used against the card's total.

Read next