Argo CD app OutOfSync right after a sync: mutating webhooks, controllers and ignoreDifferences
Synced, then OutOfSync again seconds later - something rewrites a field git declares. Find it with argocd app diff, then fix git or ignore narrowly.
Blog · topic
Pods, nodes, probes and shutdowns: Kubernetes incidents with the kubectl commands that solve them.
13 postsRSS feed
Synced, then OutOfSync again seconds later - something rewrites a field git declares. Find it with argocd app diff, then fix git or ignore narrowly.
When every app fails name lookups at the same minute, one shared resolver is down. Find it with resolv.conf, getent vs dig, and one query per hop.
137 = 128 + SIGKILL. How to tell a cgroup OOM kill from a host OOM, why -Xmx equal to the memory limit gets your Java pod killed, and how to size it.
A kubeadm cluster a year old and never upgraded stops answering kubectl. Check expiry, renew, restart the static pods, refresh your kubeconfig.
The kubelet refuses to run with swap enabled by default. How to read it from systemctl and journalctl, why swapoff alone comes back after a reboot, and the full fix.
CrashLoopBackOff is a symptom, not a cause. Read Last State and the exit code, get the logs of the previous container, check probes and events - a short, ordered checklist.
New pods stuck in CreateContainerConfigError or ContainerCreating after a config cleanup? Read the Events, find the missing ConfigMap or key, fix the config.
SIGTERM, the grace period, SIGKILL - and the endpoint removal running at the same time. Why pods drop requests on deploy, and how preStop and graceful shutdown fix it.
"0/3 nodes are available" is a per-node tally. Decode each reason - Insufficient cpu, untolerated taint, node affinity, unbound PVC - and fix the right one.
A deleted PVC that stays Terminating is protected, not broken. Find the pod that still mounts it, and why removing the finalizer is the wrong fix.
20-40 seconds of 503s on every rollout of a two-second app. Check the strategy, then readiness, then shutdown - and prove zero downtime with one rollout.
Healthy pods, a Service, and Connection refused. Check the selector (empty endpoints), readiness (ready false) and targetPort - in that order.
The grey router page means the router could not hand the request to a pod. Wrong host or path, no ready endpoints, or a Route pointing at the wrong port - and the commands that tell them apart.