ComfyUI out of memory: how to fix it on a rented GPU

ComfyUI runs out of memory when a workflow needs more VRAM than the GPU has. On your own PC you squeeze the workflow to fit. On a rented GPU the quickest fix is often a bigger card for an hour: a 48 GB RTX 6000 Ada cost about $0.10 an hour more than a 24 GB RTX 4090 on RunPod’s secure cloud in October 2026. Our RunPod guide covers deploying one.

Affiliate note: RunPod and Fanvue links on this page are referral links. If you sign up through them, AI Empire may earn credits or a commission. You pay the same. How this works.

On this page
  1. What does “out of memory” mean in ComfyUI?
  2. What should you try first?
  3. When should you rent a bigger GPU instead?
  4. Why does the same workflow fit on one pod and not another?
  5. Do ComfyUI’s low-VRAM flags help?
  6. Why does ComfyUI say “Reconnecting”?
  7. What can go wrong?

What does “out of memory” mean in ComfyUI?

The GPU ran out of VRAM while ComfyUI was loading a model or working on an image. You’ll see “CUDA out of memory”, “OutOfMemoryError” or an “allocation on device” message, and the run stops. The node that failed tells you where: a loader, the sampler, or the decode at the end.

Write down three things before you fix anything:

  • Which node failed. A loader node failing means the models don’t fit. KSampler failing means the image or batch is too big for what’s left. VAE Decode failing at the very end means the final image is too big to decode in one go.
  • Which GPU the pod has. Check the pod page. People often deploy a smaller card than they meant to when their first choice was unavailable.
  • What changed since it last worked. A new size, a batch size, a second model, a second workflow running.

What should you try first?

Make the job smaller, then restart. Set batch size to 1, put the image size back to what the model is meant for, run one heavy workflow at a time, and restart ComfyUI so it starts with empty VRAM. These cost nothing and take two minutes, so try them before you pay for anything.

  1. Batch size 1. In our face workflow it’s the batch_size field in the image-size node. Queue more runs instead of a bigger batch.
  2. Normal resolution. Our face workflow makes 1024×1536. Going much bigger in one step multiplies the memory the sampler needs. Make the image at the normal size and upscale separately if you need to.
  3. Shorter video. For video models, fewer frames and a lower resolution (480p instead of 720p) cut memory the most.
  4. One heavy job per pod. Don’t run the Dataset Maker and a face workflow on the same card at the same time.
  5. Restart ComfyUI. A long session can leave memory in use. Restart and run the job alone.

If it still fails at batch 1 and normal size, the card is too small for the job. Change the card, not the workflow.

When should you rent a bigger GPU instead?

When the job fails at batch size 1 and the model’s normal size, or when it only runs by offloading and crawls. On a rented GPU you pay for every slow minute, so a bigger card for an hour often costs less than fighting the small one. Download your files, terminate the pod, and deploy the next tier up.

Prices are RunPod secure cloud as of October 2026. Check RunPod’s pricing page before you deploy.

Out of memory on VRAM Move up to VRAM Price per hour
RTX 4090 24 GB L40 or RTX 6000 Ada 48 GB $0.74 → $0.82 or $0.84
RTX 5090 32 GB RTX 6000 Ada 48 GB $0.99 → $0.84
RTX 6000 Ada, L40 or L40S 48 GB A100 80 GB $0.82 to $1.09 → $1.59
A100 or H100 SXM 80 GB RTX PRO 6000 96 GB $1.59 or $3.49 → $2.09

Worth knowing: a 48 GB card can cost less per hour than a 32 GB RTX 5090. But the cheapest hour isn’t always the cheapest job: in the owner’s estimates (not measured runs), the RTX 5090 was still the cheapest card per finished 50-photo dataset, at about $1.90. Pick the smallest tier where your job fits without offloading.

Sign up and deploy on RunPod. Our VRAM reference by task shows which tier each step of the pipeline needs.

Why does the same workflow fit on one pod and not another?

Because the card, the job size or the tool’s own memory mode changed. A workflow that runs on a 48 GB card can fail or crawl on a 24 GB one, and some tools switch to a lighter mode by themselves when VRAM is short. So “it works” can quietly mean “it works slowly”, or “it works at lower quality”.

The case that catches people in our pipeline is LoRA training. On a smaller card like the RTX 5090, the trainer switches to a lighter mode: it needs quantization and Low VRAM, and the LoRA comes out worse. That’s why the course trains on a 96 GB RTX PRO 6000, with Low VRAM off and no quantization.

For what each step of the pipeline needs, from her face to video, see how much VRAM ComfyUI needs, by task.

Do ComfyUI’s low-VRAM flags help?

They can, but they trade speed for memory, and on a pod you pay for that time. ComfyUI’s troubleshooting docs and model issues page list the flags below (checked 2 Oct 2026). They’re start options, so on a template you’d edit its start command, which takes longer than redeploying on a bigger card.

Flag What the docs say On a rented GPU
--lowvram Low VRAM mode, uses the CPU for the text encoder; the first thing to try Slower; may do nothing (see below)
--novram Next step if --lowvram isn’t enough Much slower
--cpu Very slow, works on any hardware, absolute last resort Pointless: you’re paying for a GPU you aren’t using
--disable-smart-memory Disables smart memory management Slower; ComfyUI’s code help text says it offloads models to normal RAM instead of keeping them in VRAM
--reserve-vram 2 Reserves that many GB of VRAM for the OS Not a fix on a pod: it gives ComfyUI less VRAM, not more
--cache-none Less RAM usage, but slower Mainly a system RAM setting, not a VRAM fix

One catch with --lowvram. ComfyUI’s source code (not the docs), in its launch options file as of 1 Oct 2026, says it has no effect when ComfyUI’s newer dynamic VRAM mode is on. ComfyUI’s startup code (same commit) turns that mode on by default on NVIDIA cards with PyTorch 2.8 or newer. So on a current install on an NVIDIA pod, the first flag the docs suggest may change nothing.

Why does ComfyUI say “Reconnecting”?

Your browser lost contact with ComfyUI on the pod. Either the ComfyUI process crashed, it restarted, or the pod itself stopped. It often shows up right after you press Run on something heavy. Open the pod’s logs and read the last lines before you refresh: they tell you which one happened.

  • An out-of-memory error in the log. Same fixes as above: smaller job, then a bigger card.
  • The process was killed or restarted. Wait for ComfyUI to come back on port 8188, refresh, and run the job alone.
  • The pod shows as stopped or has no GPU. Download what you need and deploy a new pod. The RunPod guide covers why a stopped pod can come back without a GPU.

What can go wrong?

  • The bigger GPU is unavailable. Take another card with the same VRAM: an L40 instead of an RTX 6000 Ada, an A100 instead of an H100.
  • You lose your outputs switching cards. A new pod has a fresh disk. Download images, datasets and LoRAs from JupyterLab before you terminate the old one.
  • You stopped the old pod instead of terminating it. Its volume disk keeps billing while stopped. Terminate it once your files are safe.
  • You uploaded a different model file to save memory and now it’s “missing”. The loader still points to the old file name. See fixing missing models in ComfyUI.
  • Training “worked” but the LoRA looks worse. It ran with quantization and Low VRAM on a 32 to 48 GB card. Train again on an RTX PRO 6000 (96 GB) if quality matters.
  • Batch size crept back up. Some workflows save it. Check it every time you load one.

Questions people ask

Does --lowvram fix ComfyUI out of memory errors?
Sometimes, but slowly. ComfyUI's docs list --lowvram first, then --novram, then --cpu as a last resort. ComfyUI's source code (not the docs, checked 2 Oct 2026) says --lowvram has no effect when its newer dynamic VRAM mode is on, and that mode is on by default on NVIDIA cards with PyTorch 2.8 or newer. On a rented GPU, a card with more VRAM is usually the quicker fix.
Why did a workflow that used to work start running out of memory?
Usually something changed: a bigger image size, a batch size above 1, a second workflow running on the same pod, or a smaller GPU picked at deploy. A long session can also leave memory in use, so restart ComfyUI and try again.
Is it worth paying for a bigger GPU just to fix one error?
For a few hours, yes. On RunPod's secure cloud in October 2026, a 48 GB L40 or RTX 6000 Ada cost $0.82 to $0.84 an hour against $0.74 for a 24 GB RTX 4090. A slow, offloading workflow on the smaller card can easily cost more in time.

Read next