Why Your Container Won't Start: A Debugging Guide

Updated 2026-09-23

Server rack cluster
Photo: Wikimedia Commons (Public domain)

Container errors are unusually bad at telling you what went wrong. CrashLoopBackOff describes what the scheduler is doing about your problem, not what your problem is. It means “this keeps dying so I’m waiting longer between attempts,” which is the orchestrator narrating its own patience back at you.

The good news is there are only about six real causes, and the order you check them in matters more than knowing any of them.

Start here, always

Before any of the specific cases below:

kubectl describe pod <name>
kubectl logs <name> --previous

That --previous flag is the one people miss. A crashlooping pod has already restarted, so plain kubectl logs shows you the current attempt, which is usually sitting at zero bytes because the container died before writing anything. The logs you want belong to the corpse, not the patient.

If --previous returns nothing either, the container is failing before your process starts. Skip to the image and command sections below, because your application never ran at all.

CrashLoopBackOff

The container starts, exits, and Kubernetes restarts it with increasing delay. The exit code in kubectl describe narrows it down fast:

Exit code 0: your process finished successfully and returned. Kubernetes expects long-running processes, so a script that completes looks identical to a crash. If the work genuinely is finite, you want a Job, not a Deployment.

Exit code 1: an ordinary application error. The logs have your answer; this isn’t a Kubernetes problem and no amount of manifest tweaking will fix it.

Exit code 137: killed with SIGKILL, almost always the OOM killer. See below.

Exit code 139: segfault. Frequently an architecture mismatch: an image built on an Apple Silicon laptop, pushed without specifying a platform, then pulled onto x86 nodes. Check the image’s architecture before you debug the application.

Exit codes 126 and 127: “cannot execute” and “not found”. Your entrypoint path is wrong, or the script lacks the executable bit, or it has Windows line endings and the shebang is being read as /bin/sh\r.

That last one costs people entire afternoons. If a script runs on your machine and returns 127 in the container, check the line endings before anything else.

ImagePullBackOff and ErrImagePull

Kubernetes cannot retrieve the image. kubectl describe pod states the reason outright, and it’s nearly always one of three things.

The tag doesn’t exist. Usually a typo, or :latest on a registry where nobody has pushed latest. Worth noting that :latest is not special: it’s a string like any other, and if nothing was tagged with it, it simply isn’t there.

The registry needs credentials you haven’t supplied. You need an imagePullSecrets entry referencing a secret of type kubernetes.io/dockerconfigjson. A generic secret containing the same JSON will not work, which is an irritating distinction to discover at 2am.

You’ve hit a rate limit. Docker Hub limits anonymous pulls per IP address, and every node behind one NAT gateway shares that address. A cluster that pulled fine yesterday can start failing today because a colleague ran a CI job. Authenticating, even with a free account, raises the ceiling substantially.

OOMKilled

Your container exceeded its memory limit and the kernel killed it. Two causes, and they need opposite fixes.

Either the limit is genuinely too low, or your application doesn’t know the limit exists. The second is more common and much more confusing, because the fix isn’t “add more memory.”

JVMs and Node processes that predate container-aware defaults read the host’s total memory and size their heaps against it. A JVM on a 64 GB node will happily plan for a heap far larger than the 512 MB you allotted, then get killed reaching for it. Modern JVMs respect cgroup limits, but plenty of images still don’t. Set -XX:MaxRAMPercentage or the equivalent explicitly rather than trusting the default.

The tell is a container that dies under load at a memory figure nowhere near what you configured.

Pending, and nothing is happening

The pod hasn’t been scheduled. kubectl describe pod lists the reason under Events, and it’s worth reading the whole thing rather than the first line.

Insufficient CPU or memory means no node has room for what you requested. Note requested, not used. A request is a reservation; asking for 4 CPUs on a 2-CPU node schedules never, no matter how idle the cluster looks.

An unbound PersistentVolumeClaim means your storage class can’t provision. On a homelab cluster this is usually because there is no default storage class at all, so the claim waits forever for something that will never arrive.

Taints and node selectors mean you’ve told the scheduler where the pod may run and nowhere matches. Single-node clusters hit this constantly, because control-plane nodes carry a NoSchedule taint by default and that’s the only node you have.

Docker build failures

Different layer of the stack, same afternoon.

A build that works locally and fails in CI is usually a cache difference. Your machine has layers from previous builds; CI starts clean, so a step that silently depended on a cached artifact now runs for real and fails. Building with --no-cache locally reproduces it.

COPY failing on a file that plainly exists means the file is outside the build context, or .dockerignore excludes it. The build context is the directory you passed to docker build, not your shell’s working directory, and COPY ../thing can never work.

If the image builds but the container exits immediately, check whether your entrypoint is a shell form or exec form. CMD npm start runs through a shell that becomes PID 1 and doesn’t forward signals, so the container ignores SIGTERM and takes the full grace period to die on every deploy. CMD ["npm", "start"] doesn’t have that problem.

The order to check things in

Most container debugging goes wrong because people start with the most interesting hypothesis instead of the most likely one. Roughly:

  1. kubectl describe pod: read the Events at the bottom first
  2. kubectl logs --previous: the dead container’s output, not the live one
  3. Exit code: it eliminates entire categories in one step
  4. Image architecture and tag: before suspecting your application
  5. Requests versus actual node capacity: for anything stuck Pending
  6. Only then, your code

Steps one and two resolve most of it. The habit worth building is reading the full Events list rather than the summary line, because Kubernetes usually did say what was wrong. It just said it four lines further down than anyone looks.