Exit code 137, OOMKilled, and the JVM that fits its heap but not its container
137 = 128 + SIGKILL. How to tell a cgroup OOM kill from a host OOM, why -Xmx equal to the memory limit gets your Java pod killed, and how to size it.
The OnCallReady blog
Production gotchas in Linux, Kubernetes and the cloud: what the output means, why it happens and the fix, in a few minutes each. And every one is a lab you can run in a real terminal.
RSS feed27 posts · free, no ads, no tracking
137 = 128 + SIGKILL. How to tell a cgroup OOM kill from a host OOM, why -Xmx equal to the memory limit gets your Java pod killed, and how to size it.
Synced, then OutOfSync again seconds later - something rewrites a field git declares. Find it with argocd app diff, then fix git or ignore narrowly.
rm removes a name, not the data. Why df and du disagree after deleting an open file, how lsof +L1 finds it, and how to get the space back without a reboot.
When every app fails name lookups at the same minute, one shared resolver is down. Find it with resolv.conf, getent vs dig, and one query per hop.
A container stuck in Restarting (1) is a crash loop. Read the exit code, the restart count and the logs of every attempt, then fix the startup, not the policy.
A full Docker host is usually container logs, which docker system df does not count. Find the space with df and du, prune safely, and set log rotation.
Linux load average counts processes waiting on disk or NFS, not just CPU. How to spot D-state processes, why kill -9 cannot touch them, and what to fix instead.
A kubeadm cluster a year old and never upgraded stops answering kubectl. Check expiry, renew, restart the static pods, refresh your kubeconfig.
The kubelet refuses to run with swap enabled by default. How to read it from systemctl and journalctl, why swapoff alone comes back after a reboot, and the full fix.
CrashLoopBackOff is a symptom, not a cause. Read Last State and the exit code, get the logs of the previous container, check probes and events - a short, ordered checklist.
New pods stuck in CreateContainerConfigError or ContainerCreating after a config cleanup? Read the Events, find the missing ConfigMap or key, fix the config.
SIGTERM, the grace period, SIGKILL - and the endpoint removal running at the same time. Why pods drop requests on deploy, and how preStop and graceful shutdown fix it.
"0/3 nodes are available" is a per-node tally. Decode each reason - Insufficient cpu, untolerated taint, node affinity, unbound PVC - and fix the right one.
A deleted PVC that stays Terminating is protected, not broken. Find the pod that still mounts it, and why removing the finalizer is the wrong fix.
20-40 seconds of 503s on every rollout of a two-second app. Check the strategy, then readiness, then shutdown - and prove zero downtime with one rollout.
Healthy pods, a Service, and Connection refused. Check the selector (empty endpoints), readiness (ready false) and targetPort - in that order.
A 502 is nginx saying its own connection to the upstream failed. How to read error.log, tell refused from timed out, and test from where nginx stands.
ENOSPC with a half-empty disk. Check inodes (df -i), then deleted-but-open files (lsof +L1), then ext4 reserved blocks (tune2fs) - what each looks like and how to fix it.
The grey router page means the router could not hand the request to a pod. Wrong host or path, no ready endpoints, or a Route pointing at the wrong port - and the commands that tell them apart.
Slow requests that end in 200 are successes, and abandoned ones are never counted. How to read p99 from Prometheus buckets and find the wait behind it.
set -euo pipefail is the right header, but it has holes - pipelines, conditions, && chains and local x=$(cmd). What each one misses and how to close it.
StartLimitIntervalSec and StartLimitBurst belong in [Unit]. In [Service] systemd ignores them with a warning you only see in systemd-analyze verify. Plus the default that never trips.
Removing one item from a count list re-indexes everything after it, so Terraform replaces them. Switch to for_each and move state with moved blocks.
A dev plan that renames prod resources is reading prod's state. How to read the plan, find the backend key, and re-init with -reconfigure, not -migrate-state.
sensitive = true only hides values in plan output. How a prod password reaches a CI log, how to confirm it without printing it, and how to rotate it.
A stale Terraform state lock blocks every pipeline run. Read Lock Info, prove the holder is dead, force-unlock that ID, then import what it left behind.
A zombie is a child that has exited but was never reaped by its parent. Why kill -9 does nothing, how to find the parent, and why containers need a PID 1 that reaps.