DNS resolution failed for every service: resolv.conf, systemd-resolved, CoreDNS
When every app fails name lookups at the same minute, one shared resolver is down. Find it with resolv.conf, getent vs dig, and one query per hop.
Blog · topic
Incident response, SLOs and on-call habits: the practice behind the commands.
11 postsRSS feed
When every app fails name lookups at the same minute, one shared resolver is down. Find it with resolv.conf, getent vs dig, and one query per hop.
A container stuck in Restarting (1) is a crash loop. Read the exit code, the restart count and the logs of every attempt, then fix the startup, not the policy.
A full Docker host is usually container logs, which docker system df does not count. Find the space with df and du, prune safely, and set log rotation.
CrashLoopBackOff is a symptom, not a cause. Read Last State and the exit code, get the logs of the previous container, check probes and events - a short, ordered checklist.
20-40 seconds of 503s on every rollout of a two-second app. Check the strategy, then readiness, then shutdown - and prove zero downtime with one rollout.
A 502 is nginx saying its own connection to the upstream failed. How to read error.log, tell refused from timed out, and test from where nginx stands.
Slow requests that end in 200 are successes, and abandoned ones are never counted. How to read p99 from Prometheus buckets and find the wait behind it.
Removing one item from a count list re-indexes everything after it, so Terraform replaces them. Switch to for_each and move state with moved blocks.
A dev plan that renames prod resources is reading prod's state. How to read the plan, find the backend key, and re-init with -reconfigure, not -migrate-state.
sensitive = true only hides values in plan output. How a prod password reaches a CI log, how to confirm it without printing it, and how to rotate it.
A stale Terraform state lock blocks every pipeline run. Read Lock Info, prove the holder is dead, force-unlock that ID, then import what it left behind.