Why Your Container Won't Start: A Debugging Guide
Updated 2026-09-23

Container errors are unusually bad at telling you what went wrong.
CrashLoopBackOff describes what the scheduler is doing about your problem,
not what your problem is. It means “this keeps dying so I’m waiting longer
between attempts,” which is the orchestrator narrating its own patience back
at you.
The good news is there are only about six real causes, and the order you check them in matters more than knowing any of them.
Start here, always
Before any of the specific cases below:
kubectl describe pod <name>
kubectl logs <name> --previous
That --previous flag is the one people miss. A crashlooping pod has
already restarted, so plain kubectl logs shows you the current attempt,
which is usually sitting at zero bytes because the container died before
writing anything. The logs you want belong to the corpse, not the patient.
If --previous returns nothing either, the container is failing before your
process starts. Skip to the image and command sections below, because your
application never ran at all.
CrashLoopBackOff
The container starts, exits, and Kubernetes restarts it with increasing
delay. The exit code in kubectl describe narrows it down fast:
Exit code 0: your process finished successfully and returned. Kubernetes expects long-running processes, so a script that completes looks identical to a crash. If the work genuinely is finite, you want a Job, not a Deployment.
Exit code 1: an ordinary application error. The logs have your answer; this isn’t a Kubernetes problem and no amount of manifest tweaking will fix it.
Exit code 137: killed with SIGKILL, almost always the OOM killer. See below.
Exit code 139: segfault. Frequently an architecture mismatch: an image built on an Apple Silicon laptop, pushed without specifying a platform, then pulled onto x86 nodes. Check the image’s architecture before you debug the application.
Exit codes 126 and 127: “cannot execute” and “not found”. Your
entrypoint path is wrong, or the script lacks the executable bit, or it has
Windows line endings and the shebang is being read as /bin/sh\r.
That last one costs people entire afternoons. If a script runs on your machine and returns 127 in the container, check the line endings before anything else.
ImagePullBackOff and ErrImagePull
Kubernetes cannot retrieve the image. kubectl describe pod states the reason
outright, and it’s nearly always one of three things.
The tag doesn’t exist. Usually a typo, or :latest on a registry where
nobody has pushed latest. Worth noting that :latest is not special: it’s
a string like any other, and if nothing was tagged with it, it simply isn’t
there.
The registry needs credentials you haven’t supplied. You need an
imagePullSecrets entry referencing a secret of type
kubernetes.io/dockerconfigjson. A generic secret containing the same JSON
will not work, which is an irritating distinction to discover at 2am.
You’ve hit a rate limit. Docker Hub limits anonymous pulls per IP address, and every node behind one NAT gateway shares that address. A cluster that pulled fine yesterday can start failing today because a colleague ran a CI job. Authenticating, even with a free account, raises the ceiling substantially.
OOMKilled
Your container exceeded its memory limit and the kernel killed it. Two causes, and they need opposite fixes.
Either the limit is genuinely too low, or your application doesn’t know the limit exists. The second is more common and much more confusing, because the fix isn’t “add more memory.”
JVMs and Node processes that predate container-aware defaults read the
host’s total memory and size their heaps against it. A JVM on a 64 GB node
will happily plan for a heap far larger than the 512 MB you allotted, then
get killed reaching for it. Modern JVMs respect cgroup limits, but plenty of
images still don’t. Set -XX:MaxRAMPercentage or the equivalent explicitly
rather than trusting the default.
The tell is a container that dies under load at a memory figure nowhere near what you configured.
Pending, and nothing is happening
The pod hasn’t been scheduled. kubectl describe pod lists the reason under
Events, and it’s worth reading the whole thing rather than the first line.
Insufficient CPU or memory means no node has room for what you requested. Note requested, not used. A request is a reservation; asking for 4 CPUs on a 2-CPU node schedules never, no matter how idle the cluster looks.
An unbound PersistentVolumeClaim means your storage class can’t provision. On a homelab cluster this is usually because there is no default storage class at all, so the claim waits forever for something that will never arrive.
Taints and node selectors mean you’ve told the scheduler where the pod may
run and nowhere matches. Single-node clusters hit this constantly, because
control-plane nodes carry a NoSchedule taint by default and that’s the
only node you have.
Docker build failures
Different layer of the stack, same afternoon.
A build that works locally and fails in CI is usually a cache difference.
Your machine has layers from previous builds; CI starts clean, so a step
that silently depended on a cached artifact now runs for real and fails.
Building with --no-cache locally reproduces it.
COPY failing on a file that plainly exists means the file is outside the build
context, or .dockerignore excludes it. The build context is the directory you
passed to docker build, not your shell’s working directory, and COPY ../thing
can never work.
If the image builds but the container exits immediately, check whether your
entrypoint is a shell form or exec form. CMD npm start runs through a
shell that becomes PID 1 and doesn’t forward signals, so the container
ignores SIGTERM and takes the full grace period to die on every deploy.
CMD ["npm", "start"] doesn’t have that problem.
The order to check things in
Most container debugging goes wrong because people start with the most interesting hypothesis instead of the most likely one. Roughly:
kubectl describe pod: read the Events at the bottom firstkubectl logs --previous: the dead container’s output, not the live one- Exit code: it eliminates entire categories in one step
- Image architecture and tag: before suspecting your application
- Requests versus actual node capacity: for anything stuck Pending
- Only then, your code
Steps one and two resolve most of it. The habit worth building is reading the full Events list rather than the summary line, because Kubernetes usually did say what was wrong. It just said it four lines further down than anyone looks.
☕ Coffee Corner
Compiling, deploying, or waiting on a render? Here's what to brew while you wait.
🏠 Smart Home Picks
Same hobbyist care applied to your network and your front door.




