<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>OnCallReady blog</title><link>https://oncallready.dev/blog/</link><atom:link href="https://oncallready.dev/blog/rss.xml" rel="self" type="application/rss+xml"/><description>Real incidents and Linux / Kubernetes / SRE gotchas, explained.</description><language>en</language><lastBuildDate>Sat, 03 Oct 2026 08:00:00 GMT</lastBuildDate>
<item><title>Argo CD app OutOfSync right after a sync: mutating webhooks, controllers and ignoreDifferences</title><link>https://oncallready.dev/blog/argocd-outofsync-ignoredifferences/</link><guid isPermaLink="true">https://oncallready.dev/blog/argocd-outofsync-ignoredifferences/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>Synced, then OutOfSync again seconds later - something rewrites a field git declares. Find it with argocd app diff, then fix git or ignore narrowly.</description><category>GitOps</category><category>Kubernetes</category></item>
<item><title>I deleted a 40 GB log and df didn't change. Where did the space go?</title><link>https://oncallready.dev/blog/deleted-log-df-not-freed/</link><guid isPermaLink="true">https://oncallready.dev/blog/deleted-log-df-not-freed/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>rm removes a name, not the data. Why df and du disagree after deleting an open file, how lsof +L1 finds it, and how to get the space back without a reboot.</description><category>Linux</category><category>Filesystem</category></item>
<item><title>DNS resolution failed for every service: resolv.conf, systemd-resolved, CoreDNS</title><link>https://oncallready.dev/blog/dns-resolution-failed-systemd-resolved/</link><guid isPermaLink="true">https://oncallready.dev/blog/dns-resolution-failed-systemd-resolved/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>When every app fails name lookups at the same minute, one shared resolver is down. Find it with resolv.conf, getent vs dig, and one query per hop.</description><category>Networking</category><category>Linux</category><category>Kubernetes</category><category>SRE</category></item>
<item><title>Docker container keeps restarting: exit codes, restart policies, docker logs</title><link>https://oncallready.dev/blog/docker-container-restarting-exit-code/</link><guid isPermaLink="true">https://oncallready.dev/blog/docker-container-restarting-exit-code/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>A container stuck in Restarting (1) is a crash loop. Read the exit code, the restart count and the logs of every attempt, then fix the startup, not the policy.</description><category>Docker</category><category>Linux</category><category>SRE</category></item>
<item><title>Docker host out of disk but docker system df looks fine</title><link>https://oncallready.dev/blog/docker-no-space-left-on-device/</link><guid isPermaLink="true">https://oncallready.dev/blog/docker-no-space-left-on-device/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>A full Docker host is usually container logs, which docker system df does not count. Find the space with df and du, prune safely, and set log rotation.</description><category>Docker</category><category>Linux</category><category>Filesystem</category><category>SRE</category></item>
<item><title>Exit code 137, OOMKilled, and the JVM that fits its heap but not its container</title><link>https://oncallready.dev/blog/exit-code-137-oomkilled-jvm/</link><guid isPermaLink="true">https://oncallready.dev/blog/exit-code-137-oomkilled-jvm/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>137 = 128 + SIGKILL. How to tell a cgroup OOM kill from a host OOM, why -Xmx equal to the memory limit gets your Java pod killed, and how to size it.</description><category>Memory</category><category>Kubernetes</category><category>JVM</category></item>
<item><title>Load average 15 on 4 cores, but the CPU is 90% idle</title><link>https://oncallready.dev/blog/high-load-idle-cpu-d-state/</link><guid isPermaLink="true">https://oncallready.dev/blog/high-load-idle-cpu-d-state/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>Linux load average counts processes waiting on disk or NFS, not just CPU. How to spot D-state processes, why kill -9 cannot touch them, and what to fix instead.</description><category>Linux</category><category>Processes</category></item>
<item><title>kubectl: x509: certificate has expired or is not yet valid (renewing kubeadm certificates)</title><link>https://oncallready.dev/blog/kubeadm-x509-certificate-has-expired/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubeadm-x509-certificate-has-expired/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>A kubeadm cluster a year old and never upgraded stops answering kubectl. Check expiry, renew, restart the static pods, refresh your kubeconfig.</description><category>Kubernetes</category></item>
<item><title>Node NotReady: kubelet stuck in activating (auto-restart) because swap is on</title><link>https://oncallready.dev/blog/kubelet-swap-node-notready/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubelet-swap-node-notready/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>The kubelet refuses to run with swap enabled by default. How to read it from systemctl and journalctl, why swapoff alone comes back after a reboot, and the full fix.</description><category>Kubernetes</category><category>systemd</category><category>Linux</category></item>
<item><title>CrashLoopBackOff: how to find out why the pod keeps restarting</title><link>https://oncallready.dev/blog/kubernetes-crashloopbackoff-debug/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubernetes-crashloopbackoff-debug/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>CrashLoopBackOff is a symptom, not a cause. Read Last State and the exit code, get the logs of the previous container, check probes and events - a short, ordered checklist.</description><category>Kubernetes</category><category>SRE</category></item>
<item><title>CreateContainerConfigError: couldn't find key in ConfigMap (and other missing config)</title><link>https://oncallready.dev/blog/kubernetes-createcontainerconfigerror-missing-configmap-key/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubernetes-createcontainerconfigerror-missing-configmap-key/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>New pods stuck in CreateContainerConfigError or ContainerCreating after a config cleanup? Read the Events, find the missing ConfigMap or key, fix the config.</description><category>Kubernetes</category></item>
<item><title>What happens between kubectl delete pod and the container disappearing</title><link>https://oncallready.dev/blog/kubernetes-pod-deletion-sigterm-grace-sigkill/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubernetes-pod-deletion-sigterm-grace-sigkill/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>SIGTERM, the grace period, SIGKILL - and the endpoint removal running at the same time. Why pods drop requests on deploy, and how preStop and graceful shutdown fix it.</description><category>Kubernetes</category><category>Processes</category></item>
<item><title>Pod stuck in Pending: how to read FailedScheduling (requests, taints, affinity, PVCs)</title><link>https://oncallready.dev/blog/kubernetes-pod-pending-failedscheduling/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubernetes-pod-pending-failedscheduling/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>&quot;0/3 nodes are available&quot; is a per-node tally. Decode each reason - Insufficient cpu, untolerated taint, node affinity, unbound PVC - and fix the right one.</description><category>Kubernetes</category></item>
<item><title>PVC stuck in Terminating: the pvc-protection finalizer and what still uses it</title><link>https://oncallready.dev/blog/kubernetes-pvc-stuck-terminating-finalizer/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubernetes-pvc-stuck-terminating-finalizer/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>A deleted PVC that stays Terminating is protected, not broken. Find the pod that still mounts it, and why removing the finalizer is the wrong fix.</description><category>Kubernetes</category></item>
<item><title>Downtime on every Kubernetes deploy: Recreate, readiness probes and the preStop race</title><link>https://oncallready.dev/blog/kubernetes-rolling-update-downtime-readiness-prestop/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubernetes-rolling-update-downtime-readiness-prestop/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>20-40 seconds of 503s on every rollout of a two-second app. Check the strategy, then readiness, then shutdown - and prove zero downtime with one rollout.</description><category>Kubernetes</category><category>SRE</category></item>
<item><title>Kubernetes Service not working: pods Running, zero restarts, no traffic</title><link>https://oncallready.dev/blog/kubernetes-service-no-endpoints-selector/</link><guid isPermaLink="true">https://oncallready.dev/blog/kubernetes-service-no-endpoints-selector/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>Healthy pods, a Service, and Connection refused. Check the selector (empty endpoints), readiness (ready false) and targetPort - in that order.</description><category>Kubernetes</category><category>Networking</category></item>
<item><title>nginx 502 Bad Gateway but the API is up: reading error.log</title><link>https://oncallready.dev/blog/nginx-502-bad-gateway-upstream/</link><guid isPermaLink="true">https://oncallready.dev/blog/nginx-502-bad-gateway-upstream/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>A 502 is nginx saying its own connection to the upstream failed. How to read error.log, tell refused from timed out, and test from where nginx stands.</description><category>Networking</category><category>Linux</category><category>Docker</category><category>SRE</category></item>
<item><title>&quot;No space left on device&quot; but df shows free space: the 3 causes, in the order to check</title><link>https://oncallready.dev/blog/no-space-left-df-shows-free/</link><guid isPermaLink="true">https://oncallready.dev/blog/no-space-left-df-shows-free/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>ENOSPC with a half-empty disk. Check inodes (df -i), then deleted-but-open files (lsof +L1), then ext4 reserved blocks (tune2fs) - what each looks like and how to fix it.</description><category>Linux</category><category>Filesystem</category></item>
<item><title>OpenShift &quot;Application is not available&quot;: the 503 and its three causes</title><link>https://oncallready.dev/blog/openshift-application-is-not-available-503/</link><guid isPermaLink="true">https://oncallready.dev/blog/openshift-application-is-not-available-503/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>The grey router page means the router could not hand the request to a pod. Wrong host or path, no ready endpoints, or a Route pointing at the wrong port - and the commands that tell them apart.</description><category>OpenShift</category><category>Kubernetes</category><category>Networking</category></item>
<item><title>p99 latency is 25 seconds and the error rate is zero: histograms, histogram_quantile and timeouts</title><link>https://oncallready.dev/blog/prometheus-p99-latency-histogram-quantile/</link><guid isPermaLink="true">https://oncallready.dev/blog/prometheus-p99-latency-histogram-quantile/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>Slow requests that end in 200 are successes, and abandoned ones are never counted. How to read p99 from Prometheus buckets and find the wait behind it.</description><category>Observability</category><category>JVM</category><category>SRE</category></item>
<item><title>Why set -e doesn't catch a failure in foo | bar (and the other gaps)</title><link>https://oncallready.dev/blog/set-e-does-not-catch-pipe/</link><guid isPermaLink="true">https://oncallready.dev/blog/set-e-does-not-catch-pipe/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>set -euo pipefail is the right header, but it has holes - pipelines, conditions, &amp;&amp; chains and local x=$(cmd). What each one misses and how to close it.</description><category>Bash</category><category>Linux</category></item>
<item><title>systemd: your StartLimitIntervalSec is in the wrong section (and silently ignored)</title><link>https://oncallready.dev/blog/systemd-startlimitintervalsec-ignored/</link><guid isPermaLink="true">https://oncallready.dev/blog/systemd-startlimitintervalsec-ignored/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>StartLimitIntervalSec and StartLimitBurst belong in [Unit]. In [Service] systemd ignores them with a warning you only see in systemd-analyze verify. Plus the default that never trips.</description><category>systemd</category><category>Linux</category></item>
<item><title>Terraform count vs for_each: removing one item destroys the rest of the list</title><link>https://oncallready.dev/blog/terraform-count-index-shift-destroy/</link><guid isPermaLink="true">https://oncallready.dev/blog/terraform-count-index-shift-destroy/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>Removing one item from a count list re-indexes everything after it, so Terraform replaces them. Switch to for_each and move state with moved blocks.</description><category>Terraform</category><category>Azure</category><category>SRE</category></item>
<item><title>Terraform plan in dev wants to replace production: wrong state, backend key or workspace</title><link>https://oncallready.dev/blog/terraform-plan-replace-production-wrong-state/</link><guid isPermaLink="true">https://oncallready.dev/blog/terraform-plan-replace-production-wrong-state/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>A dev plan that renames prod resources is reading prod's state. How to read the plan, find the backend key, and re-init with -reconfigure, not -migrate-state.</description><category>Terraform</category><category>Azure</category><category>CI/CD</category><category>SRE</category></item>
<item><title>Terraform secret leaked in a CI log: sensitive = true, TF_LOG and rotating first</title><link>https://oncallready.dev/blog/terraform-secret-leaked-in-ci-log/</link><guid isPermaLink="true">https://oncallready.dev/blog/terraform-secret-leaked-in-ci-log/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>sensitive = true only hides values in plan output. How a prod password reaches a CI log, how to confirm it without printing it, and how to rotate it.</description><category>Terraform</category><category>CI/CD</category><category>Azure</category><category>SRE</category></item>
<item><title>Terraform &quot;Error acquiring the state lock&quot;: when force-unlock is safe</title><link>https://oncallready.dev/blog/terraform-state-lock-force-unlock/</link><guid isPermaLink="true">https://oncallready.dev/blog/terraform-state-lock-force-unlock/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>A stale Terraform state lock blocks every pipeline run. Read Lock Info, prove the holder is dead, force-unlock that ID, then import what it left behind.</description><category>Terraform</category><category>Azure</category><category>CI/CD</category><category>SRE</category></item>
<item><title>Zombie processes (&lt;defunct&gt;): why kill doesn't work and what does</title><link>https://oncallready.dev/blog/zombie-processes-defunct/</link><guid isPermaLink="true">https://oncallready.dev/blog/zombie-processes-defunct/</guid><pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate><description>A zombie is a child that has exited but was never reaped by its parent. Why kill -9 does nothing, how to find the parent, and why containers need a PID 1 that reaps.</description><category>Processes</category><category>Linux</category><category>Docker</category></item>
</channel></rss>
