The shape
The problem. A health checker run by cron hangs forever on the one service that is down, prints nothing useful, and always exits 0, so nobody notices it failed. Every line of a proper ops script prevents one of those.
What you need to know already: venvs and pip (21.29), exit codes and set -e (6.1-6.5), systemd timers (2.16), timeouts and retries (21.18, 21.22), HTTP (9.21).
Read the script once top to bottom; each part is explained below. Python basics you will see: import loads a module, def defines a function, indentation (not braces) marks blocks, class X(Exception) defines a new exception type, @dataclass makes a simple record class, and name: str / -> int are type hints (like TypeScript types, but not enforced).
#!/usr/bin/env python3
"""Check that services answer their health endpoint."""
import argparse
import json
import logging
import subprocess
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
log = logging.getLogger("check")
class CheckError(Exception):
"""A failure we expect and report, as opposed to a bug."""
@dataclass
class Result:
name: str
url: str
ok: bool
detail: str = ""
def session() -> requests.Session:
s = requests.Session()
retry = Retry(total=2, backoff_factor=0.5, status_forcelist=[502, 503, 504])
s.mount("http://", HTTPAdapter(max_retries=retry))
s.mount("https://", HTTPAdapter(max_retries=retry))
return s
def unit_active(unit: str) -> bool:
r = subprocess.run(["systemctl", "is-active", unit], capture_output=True, text=True, timeout=5)
return r.returncode == 0
def check(s: requests.Session, name: str, url: str, timeout: float) -> Result:
try:
r = s.get(url, timeout=(1, timeout)) # (connect, read)
r.raise_for_status()
return Result(name, url, r.json().get("status") == "UP", r.text.strip())
except requests.exceptions.RequestException as e:
return Result(name, url, False, type(e).__name__)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__)
p.add_argument("--timeout", type=float, default=2.0, help="read timeout, seconds")
p.add_argument("--out", type=Path, help="write results as JSON here")
p.add_argument("-v", "--verbose", action="store_true")
a = p.parse_args(argv)
logging.basicConfig(level=logging.DEBUG if a.verbose else logging.INFO,
format="%(asctime)s %(levelname)s %(name)s: %(message)s", stream=sys.stdout)
...
return 2 if failed else 0
if __name__ == "__main__":
sys.exit(main())
Every piece is there for a reason. The rest of the lesson is the reasons.
argparse
# check.py = the script above (you write it in the next mission)
./check.py --help
usage: check.py [-h] [--timeout TIMEOUT] [--out OUT] [-v]
...
./check.py --timeout abc; echo $?
usage: check.py [-h] [--timeout TIMEOUT] [--out OUT] [-v]
check.py: error: argument --timeout: invalid float value: 'abc'
2
type= converts and validates, choices=[...] restricts, action="store_true" makes a flag, required=True and positional arguments with nargs. Usage errors exit 2 by convention. click gives nicer help and subcommands with decorators; argparse needs nothing installed.
logging, not print
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(name)s: %(message)s")
log.info("%s ok=%s", name, ok) # lazy formatting: args are only rendered if the level is enabled
log.warning("%s: unit not active", unit)
log.exception("failed") # inside an except: logs the traceback too
2026-09-23 10:15:02,702 INFO check: orders ok=True {"status":"UP"}
2026-09-23 10:15:02,702 WARNING check: payments: unit not active
2026-09-23 10:15:03,708 ERROR check: unhealthy: payments
A tool that runs unattended (cron, a systemd timer, a CI job) needs timestamps, levels, a logger name, and a way to turn detail up (-v) without editing code. print gives you none of that. Log to stdout/stderr and let the platform collect it: journald, the container runtime, the CI log. (Default basicConfig goes to stderr; stream=sys.stdout if you prefer.) For machine parsing, a JSON formatter.
requests: timeouts, always
requests.get(url) # NO timeout: waits forever on a hung server
requests.get(url, timeout=2) # 2 s connect AND 2 s read (per socket read, not total)
requests.get(url, timeout=(1, 5)) # 1 s connect, 5 s read
requests has no default timeout. A health checker without one hangs on exactly the service it is supposed to report. The exceptions:
requests.exceptions.ConnectTimeout could not connect in time
requests.exceptions.ReadTimeout connected, server went quiet
requests.exceptions.ConnectionError refused, DNS failure, reset (ConnectTimeout is a subclass)
requests.exceptions.HTTPError raise_for_status() on 4xx/5xx
requests.exceptions.RequestException the base of all of them
A Session reuses connections (keep-alive) and carries defaults. Mount an HTTPAdapter with a urllib3 Retry for retries with backoff on the statuses and methods that are safe to retry (allowed_methods defaults to idempotent ones - GET, HEAD, PUT, DELETE, OPTIONS, TRACE - not POST).
subprocess.run
subprocess.run(["systemctl", "is-active", unit], capture_output=True, text=True, timeout=5)
subprocess.run(["kubectl", "get", "pods", "-n", ns], check=True) # raise on non-zero
- A list of arguments, never a string with
shell=Truebuilt from input:subprocess.run(f"systemctl restart {name}", shell=True)withname = "x; rm -rf /"is a shell injection. With a list there is no shell to inject into. check=TrueraisesCalledProcessErroron a non-zero exit:Command '['kubectl', 'get', 'pods', '-n', 'nope']' returned non-zero exit status 1.Without it, failures pass silently - the Python version of bash withoutset -e.capture_output=True, text=Truegives your.stdout/r.stderras strings.timeout=- the same lesson again: a command can hang too.
pathlib
from pathlib import Path
cfg = Path("/etc/orders") / "app.conf" # / joins paths
text = cfg.read_text()
out = Path(a.out); out.parent.mkdir(parents=True, exist_ok=True); out.write_text(data)
for p in Path("/var/log/app").glob("*.log"): print(p.name, p.stat().st_size)
os.path.join, open()/read()/close() and string slicing on paths all work and are all worse.
dataclasses, json, type hints
@dataclass
class Result:
name: str
ok: bool
detail: str = ""
json.dumps([asdict(r) for r in results], indent=2)
def check(url: str, timeout: float = 2.0) -> Result: ...
Type hints are not checked at runtime; they document and let an editor or mypy check. Read list[str] | None as "a list of strings, or None".
Exit codes and exceptions
0 success 1 unexpected failure (an uncaught exception exits 1)
2 usage error (argparse) 3+ your own meanings - document them
130 interrupted (Ctrl+C: KeyboardInterrupt, 128 + SIGINT)
sys.exit(main()) turns main's return value into the process exit code - which is what cron, systemd, CI and a calling script see. Raise a custom exception for failures you expect (class CheckError(Exception)), catch it at the top, log it, and return a meaningful code. Let genuine bugs raise and exit 1 with a traceback.
The Ctrl+C test
A script with no timeout, pointed at a hung service:
# check_old.py = the same check without a timeout (not on this box)
./check_old.py
^CTraceback (most recent call last):
File "/home/learner/.../check_old.py", line 9, in <module>
r = requests.get("http://localhost:8081/actuator/health")
...
KeyboardInterrupt
echo $?
130
The traceback (Python's stack trace) shows where it was stuck. Under cron nobody presses Ctrl+C - it just runs forever, and the next run starts another copy.
What you can now do
- Build a CLI with argparse, log with the logging module, and exit with meaningful codes.
- Call HTTP with requests and a timeout (always), and run commands with
subprocess.run([...], check=True). - Say why
shell=Trueon input is a shell injection.