RunPod pod stuck on initializing or not ready: what to check

A RunPod pod stuck on initializing is usually still busy: pulling its container image, then downloading models on first boot, which is about 30 GB on our Dataset Maker template. Work through 7 checks in order: logs, telemetry, first-boot downloads, ports, GPU count, the CUDA filter, then the template itself. If you haven’t deployed a pod before, start with our RunPod for AI images guide.

On this page
  1. What does each RunPod status or error actually mean?
  2. Which checks fix a stuck pod, in what order?
  3. How long should the first boot take?
  4. What if the GPU you want isn’t available?
  5. When is it the template, not RunPod?
  6. What can go wrong?

What does each RunPod status or error actually mean?

The status in the console and what’s really happening are often two different things. “Running” means the container started, not that ComfyUI or JupyterLab is ready. “Ready” on JupyterLab only means its server answered a ping. The table maps what you see to the real cause, so you fix the right thing first.

What you see What’s happening What to do
Pod sits on initializing The container image is still being pulled, or a startup command failed Open Logs. Look for download progress or an error
Running, but ComfyUI (8188) won’t open The template is still downloading models on first boot Wait for the logs to say the download finished and ComfyUI started
JupyterLab says Ready, page is blank Jupyter answered a status check but hasn’t finished loading Wait 60 s, hard refresh (Ctrl/Cmd + Shift + R), try a private window
502 on a port The app on that port isn’t running, or the pod has 0 GPUs Check the GPU count on the pod, then the logs
“Zero GPU Pods” after a restart Someone rented that machine’s GPU while your pod was stopped Start with 0 GPUs to rescue files, or terminate and redeploy
“OCI runtime create failed” Often a CUDA mismatch between the machine’s driver and the image Redeploy with the CUDA Versions filter set
“Driver too old” in the logs Same mismatch, caught by PyTorch Redeploy with the CUDA filter (13.0 for our trainer)
JupyterLab asks for a token Token login is on Run jupyter server list in the web terminal, paste the token

Sources: RunPod’s JupyterLab blank page, 502 errors, zero GPU, manage pods and token pages, checked 2 Oct 2026.

Which checks fix a stuck pod, in what order?

Go from cheapest to most drastic. Reading logs costs nothing. Waiting costs a few cents. Redeploying throws away whatever the pod already downloaded, so it comes last. On our templates you’ll often stop at step 3, because the pod was never broken: it was busy downloading models on its first boot.

  1. Read the logs. On the Pods page, expand the pod and click Logs. Container logs show the app’s own output. System logs show startup and shutdown events.
  2. Check telemetry. RunPod says the most reliable “is it ready” signal is the Telemetry tab: if the pod is sending GPU and CPU stats, it’s up. Individual services can still need a few minutes after that.
  3. Let first boot finish. If the logs show models downloading, leave it. Don’t restart: you’d only wait again.
  4. Check the ports. The port you click must be in the template’s exposed HTTP ports. Ours: 8188 and 8888 on the Dataset Maker, 8675 and 8888 on the LoRA trainer.
  5. Check the GPU count. The pod should show something like “1 x RTX 5090”. If it shows 0, every web button is dead even if it’s lit.
  6. Check for a CUDA error. “OCI runtime create failed” or “driver too old” means you need a newer host. See how to fix the CUDA driver error on RunPod.
  7. Suspect the template. After three restarts with no luck, stop restarting. Check the template’s README for extra steps, then redeploy or contact whoever made the template.

How long should the first boot take?

Longer than you think. RunPod says to wait 30 to 60 seconds after starting a pod before opening JupyterLab. A template that downloads models on first boot must finish that download before its web app answers. Our Dataset Maker pulls about 30 GB. Big video templates are slower: LTX pods have taken us 20 to 40 minutes to load.

Two things follow from that:

  • You pay while it downloads. Pods are billed per second (RunPod pricing), and the GPU is billed while the pod runs, download included. Pick your GPU before you deploy, not after. How RunPod billing works explains the rest.
  • Restarting the same pod skips the download; a new pod doesn’t. Models land on the pod’s volume disk at /workspace, and the Dataset Maker README says it downloads them on the first boot only. We haven’t timed a restart. A new pod downloads everything again. But a stopped pod bills its volume at $0.20 per GB per month (RunPod pricing, October 2026) and can come back with 0 GPUs, so a stop and restart only makes sense within the same day.

Our templates also have a 30 GB container disk and a 100 GB volume. If you shrink the volume to save a few cents, a 30 GB download plus your outputs can run out of room.

What if the GPU you want isn’t available?

Take the next GPU in the same VRAM tier for the job, not a smaller card. A smaller card might deploy faster but then runs out of memory halfway through the job. The deploy page marks cards as available or out of capacity. If you attached a network volume, you only see GPUs in that one data center.

You wanted Job tier Swap to (same tier or up)
RTX 4090 24 GB RTX 3090 (slower), or RTX 5090
RTX 5090 32 GB RTX 6000 Ada, L40 or L40S (48 GB)
RTX 6000 Ada 48 GB L40, L40S, or an 80 GB card

The 96 GB RTX PRO 6000 has no same-tier swap on our list. One exception: the trainer’s README lists an A100 or H100 (80 GB) as full quality too, with quantization off. The course itself trains on the RTX PRO 6000; 32 to 48 GB cards need quantization and Low VRAM, which gives a worse LoRA. For anything else that needs the 96 GB, wait or use Deploy when available (below).

Prices for each card are in the GPU table on our RunPod guide. If you’re unsure how much VRAM the job needs, see how much VRAM ComfyUI needs.

Network volumes narrow the list. A network volume is tied to one data center, and RunPod’s deploy page filters GPUs to that data center. Whether you need one at all is covered in do you need a network volume.

Out of capacity on the card you need? RunPod’s newer deploy page lets you pick an out-of-capacity GPU and choose Deploy when available, which queues the pod and deploys it when a card frees up. If you don’t have enough credit when the card frees up, the subscription fails and nothing deploys (RunPod docs, checked 2 Oct 2026).

Zero GPUs after a restart? Stopping releases your GPU, and someone else can rent it. You get three options: start with zero GPUs to copy your files off, wait and retry, or terminate and redeploy on any machine with that GPU free. A zero-GPU pod is reachable by terminal or Cloud Sync only; see how to move files off a pod.

When is it the template, not RunPod?

When the logs show the same error on every restart, or JupyterLab never loads after three restarts and long waits. RunPod’s own docs say to treat that as a template or configuration problem. Community templates aren’t supported by RunPod, so the fix comes from the template’s README or its creator, not RunPod support.

On our templates, the README covers the steps the template needs. The LoRA trainer also writes its own log to /workspace/aitk/aiempire.log, which is more readable than the raw container log when a training job won’t start.

What can go wrong?

  • You restart while it’s downloading. You interrupt the download and add more billed waiting. Read the logs first.
  • You edit a running pod to add a port. RunPod resets the pod when you edit it, which erases everything outside /workspace or a network volume. Fix ports in the template before deploying, not after.
  • You can’t deploy at all. RunPod needs at least one hour of credit for the GPU and disks you picked. Add credit or pick a cheaper card.
  • ComfyUI opens, then dies on the first run. That’s usually memory, not startup. See ComfyUI out of memory fixes.
  • JupyterLab opens for anyone. If you skipped JUPYTER_PASSWORD, anyone with the pod’s URL can open it. Set it before you deploy.
  • A stopped pod comes back with 0 GPUs and you need the files. Start it with zero GPUs and copy them off with Cloud Sync, or with runpodctl or SCP from the terminal. JupyterLab and ComfyUI won’t work without a GPU attached (RunPod 502 docs).

Questions people ask

How long should a RunPod pod take to start?
RunPod says to wait at least 30 to 60 seconds after starting a pod before opening JupyterLab. A template that downloads models on first boot takes much longer: our Dataset Maker pulls about 30 GB before ComfyUI opens, and LTX video pods have taken us 20 to 40 minutes to load. The Dataset Maker README says the models download on the first boot only. We haven't timed a restart.
Do I pay while the pod is initializing?
Yes. Pods are billed per second, and the GPU is yours, and billed, while the pod runs, downloads included. That's why picking the right GPU before you deploy matters more than restarting.
Should I restart a pod that's stuck?
Not straight away. Read the logs first. If they show a download in progress, restarting just makes you wait again. Restart only when the logs show an error or nothing at all, and treat three failed restarts as a template problem.
Why does JupyterLab say Ready but show a blank page?
RunPod marks JupyterLab Ready as soon as the server answers a status check, before the interface has loaded. Wait 60 seconds, hard refresh, or open the link in a private window.

Read next