RunPod pod stuck on initializing or not ready: what to check
A RunPod pod stuck on initializing is usually still busy: pulling its container image, then downloading models on first boot, which is about 30 GB on our Dataset Maker template. Work through 7 checks in order: logs, telemetry, first-boot downloads, ports, GPU count, the CUDA filter, then the template itself. If you haven’t deployed a pod before, start with our RunPod for AI images guide.
What does each RunPod status or error actually mean?
The status in the console and what’s really happening are often two different things. “Running” means the container started, not that ComfyUI or JupyterLab is ready. “Ready” on JupyterLab only means its server answered a ping. The table maps what you see to the real cause, so you fix the right thing first.
| What you see | What’s happening | What to do |
|---|---|---|
| Pod sits on initializing | The container image is still being pulled, or a startup command failed | Open Logs. Look for download progress or an error |
| Running, but ComfyUI (8188) won’t open | The template is still downloading models on first boot | Wait for the logs to say the download finished and ComfyUI started |
| JupyterLab says Ready, page is blank | Jupyter answered a status check but hasn’t finished loading | Wait 60 s, hard refresh (Ctrl/Cmd + Shift + R), try a private window |
| 502 on a port | The app on that port isn’t running, or the pod has 0 GPUs | Check the GPU count on the pod, then the logs |
| “Zero GPU Pods” after a restart | Someone rented that machine’s GPU while your pod was stopped | Start with 0 GPUs to rescue files, or terminate and redeploy |
| “OCI runtime create failed” | Often a CUDA mismatch between the machine’s driver and the image | Redeploy with the CUDA Versions filter set |
| “Driver too old” in the logs | Same mismatch, caught by PyTorch | Redeploy with the CUDA filter (13.0 for our trainer) |
| JupyterLab asks for a token | Token login is on | Run jupyter server list in the web terminal, paste the token |
Sources: RunPod’s JupyterLab blank page, 502 errors, zero GPU, manage pods and token pages, checked 2 Oct 2026.
Which checks fix a stuck pod, in what order?
Go from cheapest to most drastic. Reading logs costs nothing. Waiting costs a few cents. Redeploying throws away whatever the pod already downloaded, so it comes last. On our templates you’ll often stop at step 3, because the pod was never broken: it was busy downloading models on its first boot.
- Read the logs. On the Pods page, expand the pod and click Logs. Container logs show the app’s own output. System logs show startup and shutdown events.
- Check telemetry. RunPod says the most reliable “is it ready” signal is the Telemetry tab: if the pod is sending GPU and CPU stats, it’s up. Individual services can still need a few minutes after that.
- Let first boot finish. If the logs show models downloading, leave it. Don’t restart: you’d only wait again.
- Check the ports. The port you click must be in the template’s exposed HTTP ports. Ours: 8188 and 8888 on the Dataset Maker, 8675 and 8888 on the LoRA trainer.
- Check the GPU count. The pod should show something like “1 x RTX 5090”. If it shows 0, every web button is dead even if it’s lit.
- Check for a CUDA error. “OCI runtime create failed” or “driver too old” means you need a newer host. See how to fix the CUDA driver error on RunPod.
- Suspect the template. After three restarts with no luck, stop restarting. Check the template’s README for extra steps, then redeploy or contact whoever made the template.
How long should the first boot take?
Longer than you think. RunPod says to wait 30 to 60 seconds after starting a pod before opening JupyterLab. A template that downloads models on first boot must finish that download before its web app answers. Our Dataset Maker pulls about 30 GB. Big video templates are slower: LTX pods have taken us 20 to 40 minutes to load.
Two things follow from that:
- You pay while it downloads. Pods are billed per second (RunPod pricing), and the GPU is billed while the pod runs, download included. Pick your GPU before you deploy, not after. How RunPod billing works explains the rest.
- Restarting the same pod skips the download; a new pod doesn’t. Models land on the pod’s volume disk at
/workspace, and the Dataset Maker README says it downloads them on the first boot only. We haven’t timed a restart. A new pod downloads everything again. But a stopped pod bills its volume at $0.20 per GB per month (RunPod pricing, October 2026) and can come back with 0 GPUs, so a stop and restart only makes sense within the same day.
Our templates also have a 30 GB container disk and a 100 GB volume. If you shrink the volume to save a few cents, a 30 GB download plus your outputs can run out of room.
What if the GPU you want isn’t available?
Take the next GPU in the same VRAM tier for the job, not a smaller card. A smaller card might deploy faster but then runs out of memory halfway through the job. The deploy page marks cards as available or out of capacity. If you attached a network volume, you only see GPUs in that one data center.
| You wanted | Job tier | Swap to (same tier or up) |
|---|---|---|
| RTX 4090 | 24 GB | RTX 3090 (slower), or RTX 5090 |
| RTX 5090 | 32 GB | RTX 6000 Ada, L40 or L40S (48 GB) |
| RTX 6000 Ada | 48 GB | L40, L40S, or an 80 GB card |
The 96 GB RTX PRO 6000 has no same-tier swap on our list. One exception: the trainer’s README lists an A100 or H100 (80 GB) as full quality too, with quantization off. The course itself trains on the RTX PRO 6000; 32 to 48 GB cards need quantization and Low VRAM, which gives a worse LoRA. For anything else that needs the 96 GB, wait or use Deploy when available (below).
Prices for each card are in the GPU table on our RunPod guide. If you’re unsure how much VRAM the job needs, see how much VRAM ComfyUI needs.
Network volumes narrow the list. A network volume is tied to one data center, and RunPod’s deploy page filters GPUs to that data center. Whether you need one at all is covered in do you need a network volume.
Out of capacity on the card you need? RunPod’s newer deploy page lets you pick an out-of-capacity GPU and choose Deploy when available, which queues the pod and deploys it when a card frees up. If you don’t have enough credit when the card frees up, the subscription fails and nothing deploys (RunPod docs, checked 2 Oct 2026).
Zero GPUs after a restart? Stopping releases your GPU, and someone else can rent it. You get three options: start with zero GPUs to copy your files off, wait and retry, or terminate and redeploy on any machine with that GPU free. A zero-GPU pod is reachable by terminal or Cloud Sync only; see how to move files off a pod.
When is it the template, not RunPod?
When the logs show the same error on every restart, or JupyterLab never loads after three restarts and long waits. RunPod’s own docs say to treat that as a template or configuration problem. Community templates aren’t supported by RunPod, so the fix comes from the template’s README or its creator, not RunPod support.
On our templates, the README covers the steps the template needs. The LoRA trainer also writes its own log to /workspace/aitk/aiempire.log, which is more readable than the raw container log when a training job won’t start.
What can go wrong?
- You restart while it’s downloading. You interrupt the download and add more billed waiting. Read the logs first.
- You edit a running pod to add a port. RunPod resets the pod when you edit it, which erases everything outside
/workspaceor a network volume. Fix ports in the template before deploying, not after. - You can’t deploy at all. RunPod needs at least one hour of credit for the GPU and disks you picked. Add credit or pick a cheaper card.
- ComfyUI opens, then dies on the first run. That’s usually memory, not startup. See ComfyUI out of memory fixes.
- JupyterLab opens for anyone. If you skipped
JUPYTER_PASSWORD, anyone with the pod’s URL can open it. Set it before you deploy. - A stopped pod comes back with 0 GPUs and you need the files. Start it with zero GPUs and copy them off with Cloud Sync, or with runpodctl or SCP from the terminal. JupyterLab and ComfyUI won’t work without a GPU attached (RunPod 502 docs).
Questions people ask
How long should a RunPod pod take to start?
Do I pay while the pod is initializing?
Should I restart a pod that's stuck?
Why does JupyterLab say Ready but show a blank page?
Read next
- RunPod for AI images: templates, GPUs and what it costs
Rent a GPU by the hour instead of buying one. Which RunPod GPU to pick, what storage costs, how to run ComfyUI, and the stop vs terminate trap.
- "CUDA driver version is insufficient" on RunPod: how to fix it
The pod's image needs a newer CUDA than the machine's driver. On RunPod you redeploy with the CUDA Versions filter, e.g. 13.0 for our LoRA trainer.