OnCallReady

Lesson 21.31 · Spring Boot Runtime, Resilience & Python Ops · 18 min read

An ops script done properly: argparse, logging, requests, subprocess, exit codes

In plain words

Imagine two ways of asking a neighbour to check your house while you're away. One neighbour wanders round, shouts "looks fine!" to nobody, and if the door is stuck, just keeps pulling at it forever. The other has a checklist, writes the time and result of each check in a notebook, gives up on a stuck door after a minute, and leaves you a clear note: all good, or problem at the back door.

An ops script done properly is the second neighbour. argparse is the checklist of options, logging is the notebook with timestamps and levels, requests with timeout=(1, 5) is giving up on the stuck door, subprocess.run([...], check=True) runs commands safely, and the exit code (0, 2, 130) is the note that cron, systemd or CI reads.

The shape

The problem. A health checker run by cron hangs forever on the one service that is down, prints nothing useful, and always exits 0, so nobody notices it failed. Every line of a proper ops script prevents one of those.

What you need to know already: venvs and pip (21.29), exit codes and set -e (6.1-6.5), systemd timers (2.16), timeouts and retries (21.18, 21.22), HTTP (9.21).

Read the script once top to bottom; each part is explained below. Python basics you will see: import loads a module, def defines a function, indentation (not braces) marks blocks, class X(Exception) defines a new exception type, @dataclass makes a simple record class, and name: str / -> int are type hints (like TypeScript types, but not enforced).

#!/usr/bin/env python3
"""Check that services answer their health endpoint."""
import argparse
import json
import logging
import subprocess
import sys
from dataclasses import dataclass, asdict
from pathlib import Path

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

log = logging.getLogger("check")


class CheckError(Exception):
    """A failure we expect and report, as opposed to a bug."""


@dataclass
class Result:
    name: str
    url: str
    ok: bool
    detail: str = ""


def session() -> requests.Session:
    s = requests.Session()
    retry = Retry(total=2, backoff_factor=0.5, status_forcelist=[502, 503, 504])
    s.mount("http://", HTTPAdapter(max_retries=retry))
    s.mount("https://", HTTPAdapter(max_retries=retry))
    return s


def unit_active(unit: str) -> bool:
    r = subprocess.run(["systemctl", "is-active", unit], capture_output=True, text=True, timeout=5)
    return r.returncode == 0


def check(s: requests.Session, name: str, url: str, timeout: float) -> Result:
    try:
        r = s.get(url, timeout=(1, timeout))       # (connect, read)
        r.raise_for_status()
        return Result(name, url, r.json().get("status") == "UP", r.text.strip())
    except requests.exceptions.RequestException as e:
        return Result(name, url, False, type(e).__name__)


def main(argv: list[str] | None = None) -> int:
    p = argparse.ArgumentParser(description=__doc__)
    p.add_argument("--timeout", type=float, default=2.0, help="read timeout, seconds")
    p.add_argument("--out", type=Path, help="write results as JSON here")
    p.add_argument("-v", "--verbose", action="store_true")
    a = p.parse_args(argv)
    logging.basicConfig(level=logging.DEBUG if a.verbose else logging.INFO,
                        format="%(asctime)s %(levelname)s %(name)s: %(message)s", stream=sys.stdout)
    ...
    return 2 if failed else 0


if __name__ == "__main__":
    sys.exit(main())

Every piece is there for a reason. The rest of the lesson is the reasons.

argparse

# check.py = the script above (you write it in the next mission)
./check.py --help
usage: check.py [-h] [--timeout TIMEOUT] [--out OUT] [-v]
...
./check.py --timeout abc; echo $?
usage: check.py [-h] [--timeout TIMEOUT] [--out OUT] [-v]
check.py: error: argument --timeout: invalid float value: 'abc'
2

type= converts and validates, choices=[...] restricts, action="store_true" makes a flag, required=True and positional arguments with nargs. Usage errors exit 2 by convention. click gives nicer help and subcommands with decorators; argparse needs nothing installed.

logging, not print

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(name)s: %(message)s")
log.info("%s ok=%s", name, ok)           # lazy formatting: args are only rendered if the level is enabled
log.warning("%s: unit not active", unit)
log.exception("failed")                  # inside an except: logs the traceback too
2026-09-23 10:15:02,702 INFO check: orders ok=True {"status":"UP"}
2026-09-23 10:15:02,702 WARNING check: payments: unit not active
2026-09-23 10:15:03,708 ERROR check: unhealthy: payments

A tool that runs unattended (cron, a systemd timer, a CI job) needs timestamps, levels, a logger name, and a way to turn detail up (-v) without editing code. print gives you none of that. Log to stdout/stderr and let the platform collect it: journald, the container runtime, the CI log. (Default basicConfig goes to stderr; stream=sys.stdout if you prefer.) For machine parsing, a JSON formatter.

requests: timeouts, always

requests.get(url)                         # NO timeout: waits forever on a hung server
requests.get(url, timeout=2)              # 2 s connect AND 2 s read (per socket read, not total)
requests.get(url, timeout=(1, 5))         # 1 s connect, 5 s read

requests has no default timeout. A health checker without one hangs on exactly the service it is supposed to report. The exceptions:

requests.exceptions.ConnectTimeout     could not connect in time
requests.exceptions.ReadTimeout        connected, server went quiet
requests.exceptions.ConnectionError    refused, DNS failure, reset (ConnectTimeout is a subclass)
requests.exceptions.HTTPError          raise_for_status() on 4xx/5xx
requests.exceptions.RequestException   the base of all of them

A Session reuses connections (keep-alive) and carries defaults. Mount an HTTPAdapter with a urllib3 Retry for retries with backoff on the statuses and methods that are safe to retry (allowed_methods defaults to idempotent ones - GET, HEAD, PUT, DELETE, OPTIONS, TRACE - not POST).

subprocess.run

subprocess.run(["systemctl", "is-active", unit], capture_output=True, text=True, timeout=5)
subprocess.run(["kubectl", "get", "pods", "-n", ns], check=True)     # raise on non-zero

pathlib

from pathlib import Path
cfg = Path("/etc/orders") / "app.conf"          # / joins paths
text = cfg.read_text()
out = Path(a.out); out.parent.mkdir(parents=True, exist_ok=True); out.write_text(data)
for p in Path("/var/log/app").glob("*.log"): print(p.name, p.stat().st_size)

os.path.join, open()/read()/close() and string slicing on paths all work and are all worse.

dataclasses, json, type hints

@dataclass
class Result:
    name: str
    ok: bool
    detail: str = ""
json.dumps([asdict(r) for r in results], indent=2)
def check(url: str, timeout: float = 2.0) -> Result: ...

Type hints are not checked at runtime; they document and let an editor or mypy check. Read list[str] | None as "a list of strings, or None".

Exit codes and exceptions

0   success                   1   unexpected failure (an uncaught exception exits 1)
2   usage error (argparse)    3+  your own meanings - document them
130 interrupted (Ctrl+C: KeyboardInterrupt, 128 + SIGINT)

sys.exit(main()) turns main's return value into the process exit code - which is what cron, systemd, CI and a calling script see. Raise a custom exception for failures you expect (class CheckError(Exception)), catch it at the top, log it, and return a meaningful code. Let genuine bugs raise and exit 1 with a traceback.

The Ctrl+C test

A script with no timeout, pointed at a hung service:

# check_old.py = the same check without a timeout (not on this box)
./check_old.py
^CTraceback (most recent call last):
  File "/home/learner/.../check_old.py", line 9, in <module>
    r = requests.get("http://localhost:8081/actuator/health")
  ...
KeyboardInterrupt
echo $?
130

The traceback (Python's stack trace) shows where it was stuck. Under cron nobody presses Ctrl+C - it just runs forever, and the next run starts another copy.

What you can now do

Why it helps

Scripts written without timeouts and exit codes are one of the quieter sources of on-call pain. A health checker without a timeout hangs on exactly the service it should report, never alerts, and under cron piles up copies. A cleanup script that ignores a failing kubectl exits 0 and everyone thinks it worked. A subprocess.run(..., shell=True) with a name from input is a shell injection in your automation.

This lesson gives you the template you'll reuse for every tool: argparse, logging, a requests Session with retries and timeouts, subprocess with an argument list, pathlib and meaningful exit codes. It's also what reviewers look for in your own PRs, and a common practical task in platform interviews: "write a script that checks these endpoints and reports failures".

FAQ

Why logging instead of print?

A tool that runs unattended under cron, a systemd timer or CI needs timestamps, levels and a logger name, and a way to turn up detail with -v without editing code. logging gives you all of that and print gives you none. Log to stdout or stderr and let the platform collect it: journald, the container runtime, the CI log. log.exception() inside an except also records the traceback, and a JSON formatter makes logs parseable.

What does requests' timeout actually cover?

timeout=2 means 2 seconds to connect and 2 seconds per socket read, not 2 seconds in total. A server trickling bytes slowly can keep a request alive much longer. timeout=(1, 5) sets connect and read separately. Without a timeout, requests waits forever. For a hard total bound you need something extra, like running the call with an overall deadline, but connect and read timeouts already cover the common hung-server case.

Why pass a list to subprocess.run instead of a string?

With a list, the program and each argument are passed directly to the operating system, with no shell involved, so characters like ;, | or $() in a value are just characters. With a string and shell=True, a value like x; rm -rf / from user input or a config file becomes a command. Use lists, add check=True so failures raise CalledProcessError, and set timeout= because commands can hang too.

Which exit codes should my script use?

0 for success, 1 for an unexpected failure, which is what an uncaught exception gives you, and 2 for usage errors, which argparse uses. Numbers from 3 upwards are yours to define, for example 2 or 3 for "check failed", as long as you document them. Ctrl+C gives 130, 128 plus SIGINT. sys.exit(main()) passes main's return value to the process, which is what cron, systemd and CI act on.

When should I raise a custom exception?

For failures you expect and want to report cleanly, such as a service being down or a config file missing a key. Define something like class CheckError(Exception), raise it deep in the code, catch it once at the top in main, log a clear message and return a meaningful exit code. Let genuine bugs propagate: they exit 1 with a full traceback, which is what you want for something you didn't anticipate.

In an interview Mid

What makes an operations script production-ready?

Each piece prevents one failure of an unattended script (cron, a systemd timer, a CI job):

Also asked: Why is subprocess with shell=True dangerous? · What happens when you call requests.get without a timeout? · How should a script report failure to cron or systemd?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.