OnCallReady

Spring Boot Runtime, Resilience & Python Ops: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 21 of the course.

How do you stop one slow downstream service from taking down your service? Mid

First the mechanism: with no read timeout (many clients default to none), every request that calls the hung dependency holds a thread - and often a DB connection - forever. Tomcat's 200 threads fill up, health checks start failing, the pod is restarted, its traffic moves to the other replicas and they fill up the same way. A cascading failure from one missing number.

The defences, in order:

  1. Timeouts - connect and read on every client, set from the dependency's p99, fitting inside your own budget. A timeout turns a hang into an error you can count.
  2. Don't hold scarce resources across the call - no DB connection held while waiting on HTTP.
  3. Bulkhead - a bounded pool per dependency (max-concurrent-calls=20), so only 20 threads can be stuck on it.
  4. Circuit breaker - after enough failures, fail fast with a fallback; it needs the timeout underneath.
  5. Retries only for idempotent calls, with backoff and jitter, at one layer.
  6. Backpressure - bounded queues and a fast 503 instead of silent queueing.

Also asked: What is the difference between liveness and readiness probes, and how should a Spring Boot app use them? · Our deploys cause a burst of 502 errors every time. What would you check? · Why is a bigger connection pool often not the fix for slow responses?

What is Spring Boot Actuator and why does a platform team care about it? Mid

Actuator (spring-boot-starter-actuator) adds operational HTTP endpoints to a Spring Boot app: /actuator/health (with liveness and readiness groups), /metrics and /prometheus (numbers for monitoring), /env and /configprops (which config won), /threaddump, /heapdump, /loggers. It is how probes, monitoring and you at 3am talk to the app without a shell - the thread dump over HTTP works even when the image has no jcmd.

Only health is exposed by default; the rest is opt-in with management.endpoints.web.exposure.include=....

Why the platform cares about security: /env can show secrets (masked since Boot 3 unless show-values=always), /heapdump contains every secret in memory, /loggers lets anyone switch on DEBUG. Standard practice: management.server.port=8081, not routed through the Ingress; expose only what probes, monitoring and runbooks use.

Also asked: Actuator endpoints are reachable through the public ingress. What is the risk and how do you fix it? · How do you get a thread dump from a Spring Boot service whose image has no JDK? · Which Actuator endpoints are exposed by default, and how do you expose more?

Learn it: 21.1 Actuator: the seam between the app and the platform

What is the difference between liveness and readiness probes, and what should each check in a Spring Boot app? Mid

The groups come from management.endpoint.health.probes.enabled=true (automatic when Boot detects Kubernetes). Probes read the status code: 200 UP, 503 DOWN.

The classic mistake: the aggregate /actuator/health (which includes db) as liveness. The database blips, every pod is restarted at once, they all come back into the same dead database. A restart does not fix a database.

Also: a startup probe for slow starters (failureThreshold x periodSeconds > start time), liveness tolerant of a few failures, the default timeoutSeconds: 1 in mind, and management.server.port so probes answer even when the request threads are stuck.

Also asked: Our database went down for a minute and every pod of the service restarted. Why, and how do you prevent it? · What is a startup probe for, and how do you size it for an app that takes 90 seconds to start? · Why would you run the management endpoints on a separate port?

Learn it: 21.3 Health groups, liveness vs readiness, and the three probes

A setting you changed has no effect on the running Spring Boot app. How do you troubleshoot it? Mid

Spring Boot reads each property from many property sources and the highest one that has it wins: command-line arguments > -D system properties > environment variables > config files outside the jar > files inside the jar (profile-specific beat plain).

  1. Which source won? curl localhost:8080/actuator/env/<property> lists every source that defines it, in precedence order, with its origin. Values are masked; origins are not.
  2. Did the process get my change? systemctl show -p Environment <unit> or kubectl exec ... -- env; jcmd PID VM.command_line / VM.system_properties for -D and arguments. Env and files are read at startup - was it restarted?
  3. Is the name right? Relaxed binding: dots to underscores, dashes removed, uppercase. SPRING_CONFIG_ADDITIONAL_LOCATION is silently ignored; it is SPRING_CONFIG_ADDITIONALLOCATION.
  4. What is really used? A placeholder like ${db.pool.max:10} may read a different key; /actuator/configprops or hikaricp_connections_max shows the bound value.

Also asked: How does Spring Boot decide a property's value when it is set in several places? · How do you override a property per environment without rebuilding the image? · What is a Spring profile and when would you use one?

Learn it: 21.5 Externalised configuration and its precedence

What happens when a pod is terminated, and how do you make a Spring Boot app shut down without failing requests? Mid

When a pod is deleted, two things happen at the same time: the kubelet runs the preStop hook and then sends SIGTERM, while the control plane removes the pod from the EndpointSlices and kube-proxy and ingress controllers update - which takes a moment to reach every node. After terminationGracePeriodSeconds (default 30 s) comes SIGKILL.

What the app needs:

On a VM systemd plays the kubelet: TimeoutStopSec then SIGKILL (Failed with result 'timeout'), and a clean SIGTERM exit is 143 - hence SuccessExitStatus=143.

Also asked: Why do some services return 502s during every deployment, and what fixes it? · What is the difference between SIGTERM and SIGKILL in a container shutdown? · Why does a systemd unit for a Java service set SuccessExitStatus=143?

Learn it: 21.8 Graceful shutdown, TimeoutStopSec and terminationGracePeriodSeconds

Explain counters, gauges and timers or histograms, with an example of each from a Java service. Mid

Each label combination is its own series - which is why uri is the route template (/api/orders/{id}), not the raw path: otherwise cardinality explodes. Micrometer records them; /actuator/prometheus serves them as text.

Also asked: Which metrics would you put on a dashboard for a Spring Boot service, and why? · What is metric cardinality and why does it matter? · Why is average latency a poor signal on its own?

Learn it: 21.11 Micrometer and /actuator/prometheus

A service's latency jumped to 30 seconds but the error rate is flat and CPU is idle. What do you suspect? Mid

A full pool. A connection pool (HikariCP for the database) lends a fixed number of connections; when all are lent out, the next thread waits - up to connection-timeout, 30 s by default - and then usually gets one. So exhaustion shows as latency, not errors: the only error is Connection is not available, request timed out after 30000ms, and only after 30 s.

Confirm: hikaricp_connections_active = max and hikaricp_connections_pending > 0; a thread dump full of threads in HikariPool.getConnection.

The usual cause is holding a connection while waiting on something else - a @Transactional method that calls another service over HTTP. Little's law: 20 connections held 8 s each = 2.5 requests/s for the whole pool, for every endpoint that needs the database.

The fix is a timeout on the slow call and not holding the connection across it - not a bigger pool: the database has finite parallelism, and 10 pods x 200 connections can exceed max_connections. leak-detection-threshold shows who holds connections too long.

Also asked: What is a connection pool and why do applications use one? · Why is a bigger connection pool often slower? · How would you size a database connection pool for a service with several replicas?

Learn it: 21.15 Connection pools: why exhaustion is latency, not errors

What timeouts should an HTTP client have, and how do you choose them? Mid

Assume there is no timeout unless someone set it: HttpURLConnection and Reactor Netty have no read/response timeout by default, and on oncall-lab orders had http.client.read.timeout.ms=0 (infinite).

Choosing values - the budget: set each from the dependency's latency (a few times its p99), make the sum on the sequential path (times retries) fit inside your own SLO, propagate the deadline (pass on only what is left), and keep upstream timeouts longer than downstream ones - nginx proxy_read_timeout > the app's total > each client's read timeout.

Also asked: Explain timeout budgets across a chain of service calls. · A dependency hung and your service went down with it. Explain the mechanism. · What does a connect timeout tell you that a read timeout does not?

Learn it: 21.18 Timeouts: connect, read, total - and the budget

What is a circuit breaker and how does it work? Mid

A circuit breaker stops calling a dependency that is failing, so you fail fast (and free the thread) instead of waiting on it:

In Resilience4j it is properties per instance (sliding-window-size, minimum-number-of-calls, failure-rate-threshold). The defaults (100 calls, 60 s slow-call threshold) are slow to trip on a low-traffic service - tune them.

The catch: it needs a timeout underneath. It counts a call when the call completes; a hung call with no read timeout never completes, so the breaker never opens. Timeout first, then breaker, plus a bulkhead so only a bounded number of threads can wait on that dependency.

Also asked: When is it safe to retry a failed request? · Explain retry storms and how jitter prevents them. · What is the bulkhead pattern?

Learn it: 21.22 Retries, jitter, circuit breakers, bulkheads, backpressure

An async service has high latency, no errors and idle Tomcat threads. How do you investigate? Mid

With a CompletableFuture controller the waiting did not go away - it moved from Tomcat's threads to an executor. So I look at the executors:

Then check the code for supplyAsync without an executor: it uses ForkJoinPool.commonPool() (CPUs - 1, shared by the whole JVM), and on 2 CPUs or fewer a new thread per task. Nobody chose the concurrency.

Fixes: a named, bounded executor per dependency, a bounded queue (reject fast), a timeout inside the task (orTimeout or a read timeout), and alerts on the queue metric.

Also asked: What is the risk of CompletableFuture.supplyAsync without an executor argument? · Why is an unbounded queue in front of a thread pool dangerous? · What is the difference between synchronous and asynchronous request handling in a web service?

Learn it: 21.25 Async and CompletableFuture: when nothing throws

What is a Python virtual environment, and how would you install a small Python tool on a server? Mid

A venv is a directory with its own bin/python (a link to the system interpreter) and its own site-packages - packages for one project only, like a project-local node_modules. Nothing global changes.

Why it matters on a server: system Python belongs to apt. Tools like unattended-upgrades and cloud-init depend on exact versions of its packages, so sudo pip install can break the OS weeks later. Ubuntu blocks it (PEP 668, "externally managed environment"); --break-system-packages is named honestly.

Installing a tool:

  1. python3 -m venv /opt/tool/.venv (needs python3-venv).
  2. /opt/tool/.venv/bin/pip install -r requirements.txt - a list pinned with pip freeze (exact versions = a lockfile).
  3. Run it by the venv's interpreter path: ExecStart=/opt/tool/.venv/bin/python /opt/tool/check.py. systemd and cron do not activate anything; activate only prepends .venv/bin to PATH.

Never commit the venv; recreate it from the pinned list. For CLI tools you just want to run: pipx.

Also asked: Why is running sudo pip install on a production server dangerous? · How do you pin a Python tool's dependencies? · How would you run a Python script from a systemd timer so it uses the right packages?

Learn it: 21.29 Python for operations: the subset, venv, and why never system pip

What makes an operations script production-ready? Mid

Each piece prevents one failure of an unattended script (cron, a systemd timer, a CI job):

Also asked: Why is subprocess with shell=True dangerous? · What happens when you call requests.get without a timeout? · How should a script report failure to cron or systemd?

Learn it: 21.31 An ops script done properly: argparse, logging, requests, subprocess, exit codes

How do you test a Python ops script that calls the network or other systems? Mid

With pytest: files test_*.py, functions test_*, plain assert (pytest rewrites it to show the values on failure). Run pytest -v, -x to stop at the first failure, -k name to select.

Keep the logic in plain functions (status_to_code, main(argv) returning an exit code) so it can be called from a test, and replace the edges with fixtures:

Then assert health.main(["--url", "http://x"]) == 2 tests the exit code a cron job or a pipeline would see. pytest's own exit code (0 passed, 1 failed, 5 no tests) is what a build step relies on.

For Kubernetes: the Python client with config.load_kube_config() locally and load_incluster_config() in a pod, with a ServiceAccount Role that allows only what the tool needs.

Also asked: How should a tool running inside a pod authenticate to the Kubernetes API? · What is a pytest fixture, and which built-in ones do you use most? · How would you structure a Python ops tool so that it is easy to test?

Learn it: 21.34 pytest, fixtures, and the Kubernetes Python client

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.