OnCallReady

Lesson 21.34 · Spring Boot Runtime, Resilience & Python Ops · 10 min read

pytest, fixtures, and the Kubernetes Python client

In plain words

Imagine building a model bridge out of Lego. Before putting a real toy car on it, you test each piece: does this beam hold a small weight? Does this joint snap in? For the part that should connect to the table, you use a pretend table so you don't need the real one. If a piece breaks, you know exactly which one.

pytest is that for your Python tools. You write small functions named test_* with plain assert, and pytest shows exactly what failed, like assert 1 == 2 where 1 = status_to_code('DOWN'). Fixtures prepare things for tests, such as a temporary app.conf in tmp_path; monkeypatch swaps the real network call for a fake. The Kubernetes client library is then the real API you call, starting with load_kube_config().

pytest

The problem. An ops script that deletes pods or pages people must be right, and "I ran it once on my laptop" is not a test. pytest is Python's test runner (like Jest); with it you test the logic without a real cluster, then talk to the real one with the Kubernetes client.

What you need to know already: the ops script shape (21.31), venvs (21.29), exit codes (6.5), kubeconfig and ServiceAccounts (15.1, 17.33), pods.json from 7.11.

Files named test_*.py, functions named test_*, plain assert:

# test_health.py
from health import status_to_code

def test_up():
    assert status_to_code("UP") == 0

def test_down():
    assert status_to_code("DOWN") == 2
(.venv) $ pytest
============================= test session starts ==============================
platform linux -- Python 3.14.4, pytest-9.1.1, pluggy-1.6.0
rootdir: /home/learner/oncall-lab/labs/4a-runtime/python
collected 2 items

test_health.py .F                                                        [100%]

=================================== FAILURES ===================================
__________________________________ test_down ___________________________________

    def test_down():
>       assert status_to_code("DOWN") == 2
E       assert 1 == 2
E        +  where 1 = status_to_code('DOWN')

test_health.py:7: AssertionError
=========================== short test summary info ============================
FAILED test_health.py::test_down - assert 1 == 2
========================= 1 failed, 1 passed in 0.03s ==========================

Fixtures

A fixture is a function that prepares something; a test asks for it by naming it as a parameter (@pytest.fixture is a decorator - a label on the function, like the Java annotations in 21.15):

import pytest

@pytest.fixture
def config_file(tmp_path):
    p = tmp_path / "app.conf"
    p.write_text("db.pool.max=20\nhttp.client.read.timeout.ms=0\n")
    return p

def test_finds_missing_timeout(config_file):
    assert find_problems(config_file) == ["http.client.read.timeout.ms is 0 (no timeout)"]

Built in and worth knowing:

tmp_path       a fresh temporary directory (pathlib.Path) per test
monkeypatch    monkeypatch.setattr(module, "name", fake) / setenv / delenv - undone after the test
capsys         capture stdout/stderr: out, err = capsys.readouterr()

monkeypatch is how you test code that calls the network without a network (a lambda is Python's short anonymous function, like (url, timeout) => "DOWN"):

def test_down_service_is_exit_2(monkeypatch):
    monkeypatch.setattr(health, "fetch_status", lambda url, timeout: "DOWN")
    assert health.main(["--url", "http://x"]) == 2

Shared fixtures live in conftest.py in the test directory.

The Kubernetes client

The Kubernetes Python client (pip install kubernetes) calls the same API kubectl does, and returns Python objects instead of text:

from kubernetes import client, config

config.load_kube_config()              # ~/.kube/config (what kubectl uses)
# config.load_incluster_config()       # inside a pod: the ServiceAccount token
v1 = client.CoreV1Api()
for pod in v1.list_pod_for_all_namespaces().items:
    print(pod.metadata.namespace, pod.metadata.name, pod.status.phase)
    for c in pod.spec.containers:
        limits = c.resources.limits or {}            # None when unset
        if "memory" not in limits:
            print("  no memory limit:", c.name)

(simulator) On this box list_pod_for_all_namespaces() and list_namespaced_pod() are backed by the ~/labs/data/pods.json fixture from chapter 7, and need a ~/.kube/config file to exist. Other APIs raise a labelled NotImplementedError.

Later (Ch 22): the cloud SDKs (azure-identity, azure-mgmt-*) work the same way as the Kubernetes client, with a credential object in place of the kubeconfig.

What you can now do

Why it helps

Your scripts will make decisions: which pods to delete, which namespaces are orphaned, whether a check counts as failed. A bug in that logic, found in production, can delete the wrong thing. pytest with monkeypatch lets you test the decisions without a cluster, and gives CI a clear pass or fail through its exit code.

The client part is where your tools become useful at work: listing pods without memory limits across namespaces, finding Deployments with no readiness probe, or reporting which images run where - with the same code working from your laptop (your kubeconfig) and from inside a pod (its ServiceAccount). Those appear in real platform tickets, and in pod-based tools you'll need a ServiceAccount with a narrow Role.

FAQ

What is a fixture?

A function that prepares something a test needs, which the test requests by naming it as a parameter. @pytest.fixture def config_file(tmp_path): ... writes a file and returns its path, and def test_x(config_file): receives it. pytest calls fixtures for you, in dependency order, and cleans up afterwards. Built-in ones worth knowing are tmp_path, monkeypatch and capsys. Shared fixtures go in conftest.py.

How do I test code that calls the network without a network?

Replace the function that does the network call with monkeypatch: monkeypatch.setattr(health, "fetch_status", lambda url, timeout: "DOWN"), then call your code and assert on its result or exit code. The change is undone after the test. That keeps tests fast and deterministic, and it pushes you towards a good structure: a thin function that does I/O, and logic that takes plain values and is easy to test.

What's the difference between load_kube_config and load_incluster_config?

load_kube_config() reads ~/.kube/config, the same file kubectl uses, with your current context and credentials: right for running a tool from your laptop or a VM. load_incluster_config() is for code running inside a pod: it uses the ServiceAccount token mounted at /var/run/secrets/kubernetes.io/serviceaccount and the in-cluster API address. A tool running in a pod needs a Role granting exactly the verbs it uses, such as list pods.

How do I try both configs in one script?

Try the in-cluster config first and fall back to the kubeconfig: try: config.load_incluster_config() then except config.ConfigException: config.load_kube_config(). Inside a pod the first works (the ServiceAccount token is mounted); on your laptop it raises and the second reads ~/.kube/config. The same tool then runs in both places without a flag, and permissions follow whoever runs it: you, or the pod's ServiceAccount.

What do pytest's exit codes mean for CI?

0 means all tests passed, 1 means some failed, 2 means the run was interrupted, which includes errors while collecting tests, and 5 means no tests were collected at all. CI steps rely on the non-zero codes to fail the build. Exit code 5 is worth noticing: a renamed test file that no longer matches test_*.py makes the run "succeed" with nothing tested in some setups, so treat 5 as a failure.

In an interview Mid

How do you test a Python ops script that calls the network or other systems?

With pytest: files test_*.py, functions test_*, plain assert (pytest rewrites it to show the values on failure). Run pytest -v, -x to stop at the first failure, -k name to select.

Keep the logic in plain functions (status_to_code, main(argv) returning an exit code) so it can be called from a test, and replace the edges with fixtures:

Then assert health.main(["--url", "http://x"]) == 2 tests the exit code a cron job or a pipeline would see. pytest's own exit code (0 passed, 1 failed, 5 no tests) is what a build step relies on.

For Kubernetes: the Python client with config.load_kube_config() locally and load_incluster_config() in a pod, with a ServiceAccount Role that allows only what the tool needs.

Also asked: How should a tool running inside a pod authenticate to the Kubernetes API? · What is a pytest fixture, and which built-in ones do you use most? · How would you structure a Python ops tool so that it is easy to test?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.