OnCallReady

Chapter 21 Spring Boot Runtime, Resilience & Python Ops

Actuator and probes, config precedence, graceful shutdown, Micrometer; pools, timeouts, retries, circuit breakers and async exhaustion under real load; then the operational subset of Python.

In plain words

Think of a shop that relies on deliveries from a warehouse. If the delivery van is late and the shop assistant just stands at the door waiting, the queue at the till grows until nobody is served. A sensible shop gives up after a few minutes, tells the customer "not in stock today", and puts a sign up so staff stop checking for a while. It also has a way to tell the manager "I'm open" or "I'm closing", and a clipboard where it notes how many customers waited and for how long.

That is this chapter. Spring Boot's Actuator is the "I'm open" sign and the clipboard; timeouts, circuit breakers and bulkheads are the sensible shop rules. The Python part gives you a small, dependable way to write your own ops tools: a script that checks orders' health with a timeout and a clear exit code.

Why it matters on call

Most real outages in microservice platforms are not crashes but hangs: a slow dependency, a missing read timeout, a pool that fills up, and a service that looks healthy while doing nothing. http.client.read.timeout.ms=0 in orders' config is the textbook version. This chapter lets you read those incidents from the outside (Actuator, Micrometer metrics) and prevent them in design reviews: probes that don't restart every pod when the database blips, shutdowns that don't cause 502s on every deploy, timeouts that fit the SLO.

The Python part covers the small tools every platform engineer writes: health checkers, cleanup jobs, report generators against the Kubernetes API. It comes after the JVM chapter because it reads the same problems from the app's side, and before the cloud chapters, where you will script against real cloud APIs.

Lessons

  1. Actuator: the seam between the app and the platform
  2. Health groups, liveness vs readiness, and the three probes
  3. Externalised configuration and its precedence
  4. Graceful shutdown, TimeoutStopSec and terminationGracePeriodSeconds
  5. Micrometer and /actuator/prometheus
  6. Connection pools: why exhaustion is latency, not errors
  7. Timeouts: connect, read, total - and the budget
  8. Retries, jitter, circuit breakers, bulkheads, backpressure
  9. Async and CompletableFuture: when nothing throws
  10. Python for operations: the subset, venv, and why never system pip
  11. An ops script done properly: argparse, logging, requests, subprocess, exit codes
  12. pytest, fixtures, and the Kubernetes Python client

29 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

Why is a platform engineer learning Spring Boot internals?

Because the platform and the application meet at a few seams, and Spring Boot owns them: which endpoint the probes call, how the app shuts down on SIGTERM, how environment variables override config, which metrics the monitoring system collects. When a pod restarts in a loop, a deploy causes 502s or a config value is wrong, the answer is usually in one of those seams. You don't write the business code, but you review the Deployment YAML and the application.properties that decide how it behaves on Kubernetes.

What is the difference between resilience patterns and just scaling up?

Scaling adds capacity; resilience patterns limit the damage when something is slow or failing. If payments hangs, ten more orders pods just means ten more pods with every thread stuck waiting. A read timeout gives the thread back, a circuit breaker stops calling a dependency that is clearly down, a bulkhead limits how many threads one dependency can take, and backpressure rejects excess load quickly. Scaling helps with real load; these help with failure, and you need both.

Why Python and not just bash?

Bash is excellent for gluing commands together, but it gets fragile once you need to parse JSON, call HTTP APIs with retries, handle errors properly and test the result. Python gives you requests with timeouts, argparse, logging, real data structures and pytest, and it has an official Kubernetes client library (and SDKs for every cloud). The rule of thumb: a few commands in a row stay bash; anything with API calls, branching logic or that needs tests becomes Python.

Do I need to learn Django or pandas for platform work?

No. Platform work in Python is a small subset: venv and pinned requirements, argparse, logging, requests with timeouts, subprocess with argument lists, pathlib, json and dataclasses, custom exceptions with exit codes, pytest, and the kubernetes client library. Web frameworks, data science libraries and notebooks are other jobs. The simulator deliberately runs only that subset and says (simulator) when you step outside it.

Is Spring Boot in a Kubernetes pod configured differently than on this VM?

The mechanisms are the same, the defaults shift. On Kubernetes Boot detects KUBERNETES_SERVICE_HOST and enables the liveness and readiness health groups automatically; on the VM you set management.endpoint.health.probes.enabled=true. Config comes from env vars in the Deployment instead of a systemd drop-in, SIGTERM comes from the kubelet with terminationGracePeriodSeconds instead of systemd's TimeoutStopSec, and metrics are collected by a monitoring system calling the metrics endpoint instead of you running curl.