Spring Boot Runtime, Resilience & Python Ops: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 21 of the course.
How do you stop one slow downstream service from taking down your service? Mid
First the mechanism: with no read timeout (many clients default to none), every request that calls the hung dependency holds a thread - and often a DB connection - forever. Tomcat's 200 threads fill up, health checks start failing, the pod is restarted, its traffic moves to the other replicas and they fill up the same way. A cascading failure from one missing number.
The defences, in order:
- Timeouts - connect and read on every client, set from the dependency's p99, fitting inside your own budget. A timeout turns a hang into an error you can count.
- Don't hold scarce resources across the call - no DB connection held while waiting on HTTP.
- Bulkhead - a bounded pool per dependency (
max-concurrent-calls=20), so only 20 threads can be stuck on it. - Circuit breaker - after enough failures, fail fast with a fallback; it needs the timeout underneath.
- Retries only for idempotent calls, with backoff and jitter, at one layer.
- Backpressure - bounded queues and a fast 503 instead of silent queueing.
Also asked: What is the difference between liveness and readiness probes, and how should a Spring Boot app use them? · Our deploys cause a burst of 502 errors every time. What would you check? · Why is a bigger connection pool often not the fix for slow responses?
What is Spring Boot Actuator and why does a platform team care about it? Mid
Actuator (spring-boot-starter-actuator) adds operational HTTP endpoints to a Spring Boot app: /actuator/health (with liveness and readiness groups), /metrics and /prometheus (numbers for monitoring), /env and /configprops (which config won), /threaddump, /heapdump, /loggers. It is how probes, monitoring and you at 3am talk to the app without a shell - the thread dump over HTTP works even when the image has no jcmd.
Only health is exposed by default; the rest is opt-in with management.endpoints.web.exposure.include=....
Why the platform cares about security: /env can show secrets (masked since Boot 3 unless show-values=always), /heapdump contains every secret in memory, /loggers lets anyone switch on DEBUG. Standard practice: management.server.port=8081, not routed through the Ingress; expose only what probes, monitoring and runbooks use.
Also asked: Actuator endpoints are reachable through the public ingress. What is the risk and how do you fix it? · How do you get a thread dump from a Spring Boot service whose image has no JDK? · Which Actuator endpoints are exposed by default, and how do you expose more?
Learn it: 21.1 Actuator: the seam between the app and the platform
What is the difference between liveness and readiness probes, and what should each check in a Spring Boot app? Mid
- Liveness = "restart me". Cheap, local, fails only if a restart would fix it. In Boot:
/actuator/health/liveness=livenessStateonly. - Readiness = "stop sending me traffic". May include a dependency the app cannot work without:
management.endpoint.health.group.readiness.include=readinessState,db. It flips toOUT_OF_SERVICE(503) as soon as graceful shutdown starts.
The groups come from management.endpoint.health.probes.enabled=true (automatic when Boot detects Kubernetes). Probes read the status code: 200 UP, 503 DOWN.
The classic mistake: the aggregate /actuator/health (which includes db) as liveness. The database blips, every pod is restarted at once, they all come back into the same dead database. A restart does not fix a database.
Also: a startup probe for slow starters (failureThreshold x periodSeconds > start time), liveness tolerant of a few failures, the default timeoutSeconds: 1 in mind, and management.server.port so probes answer even when the request threads are stuck.
Also asked: Our database went down for a minute and every pod of the service restarted. Why, and how do you prevent it? · What is a startup probe for, and how do you size it for an app that takes 90 seconds to start? · Why would you run the management endpoints on a separate port?
Learn it: 21.3 Health groups, liveness vs readiness, and the three probes
A setting you changed has no effect on the running Spring Boot app. How do you troubleshoot it? Mid
Spring Boot reads each property from many property sources and the highest one that has it wins: command-line arguments > -D system properties > environment variables > config files outside the jar > files inside the jar (profile-specific beat plain).
- Which source won?
curl localhost:8080/actuator/env/<property>lists every source that defines it, in precedence order, with its origin. Values are masked; origins are not. - Did the process get my change?
systemctl show -p Environment <unit>orkubectl exec ... -- env;jcmd PID VM.command_line/VM.system_propertiesfor-Dand arguments. Env and files are read at startup - was it restarted? - Is the name right? Relaxed binding: dots to underscores, dashes removed, uppercase.
SPRING_CONFIG_ADDITIONAL_LOCATIONis silently ignored; it isSPRING_CONFIG_ADDITIONALLOCATION. - What is really used? A placeholder like
${db.pool.max:10}may read a different key;/actuator/configpropsorhikaricp_connections_maxshows the bound value.
Also asked: How does Spring Boot decide a property's value when it is set in several places? · How do you override a property per environment without rebuilding the image? · What is a Spring profile and when would you use one?
Learn it: 21.5 Externalised configuration and its precedence
What happens when a pod is terminated, and how do you make a Spring Boot app shut down without failing requests? Mid
When a pod is deleted, two things happen at the same time: the kubelet runs the preStop hook and then sends SIGTERM, while the control plane removes the pod from the EndpointSlices and kube-proxy and ingress controllers update - which takes a moment to reach every node. After terminationGracePeriodSeconds (default 30 s) comes SIGKILL.
What the app needs:
server.shutdown=graceful(default since Boot 3.4): stop accepting, finish in-flight requests, then close the DataSource. Readiness flips to 503 at once.- A preStop sleep of ~5 s, so traffic still routed to the pod is served while endpoint removal propagates. Without it: 502s on every deploy.
- The arithmetic: preStop +
spring.lifecycle.timeout-per-shutdown-phase+ the rest < the grace period. Otherwise SIGKILL mid-drain:Empty reply from server/ 502.
On a VM systemd plays the kubelet: TimeoutStopSec then SIGKILL (Failed with result 'timeout'), and a clean SIGTERM exit is 143 - hence SuccessExitStatus=143.
Also asked: Why do some services return 502s during every deployment, and what fixes it? · What is the difference between SIGTERM and SIGKILL in a container shutdown? · Why does a systemd unit for a Java service set SuccessExitStatus=143?
Learn it: 21.8 Graceful shutdown, TimeoutStopSec and terminationGracePeriodSeconds
Explain counters, gauges and timers or histograms, with an example of each from a Java service. Mid
- Counter - only goes up: requests served,
http_server_requests_seconds_count. You look at its rate of change, not its value. - Gauge - a current value that goes up and down:
hikaricp_connections_active,hikaricp_connections_pending(threads waiting for a DB connection - the early warning of pool exhaustion). - Timer - count + total time + max:
_count,_sum,_max. Average latency = sum / count; for a recent average, divide the change in sum by the change in count between two readings. - Histogram - counts per latency bucket (
_bucket{le="0.1"}, "at most 0.1 s"). Needed for percentiles: an average hides the slow tail, and p99 is computed from the buckets. In Boot:management.metrics.distribution.percentiles-histogram.http.server.requests=true.
Each label combination is its own series - which is why uri is the route template (/api/orders/{id}), not the raw path: otherwise cardinality explodes. Micrometer records them; /actuator/prometheus serves them as text.
Also asked: Which metrics would you put on a dashboard for a Spring Boot service, and why? · What is metric cardinality and why does it matter? · Why is average latency a poor signal on its own?
A service's latency jumped to 30 seconds but the error rate is flat and CPU is idle. What do you suspect? Mid
A full pool. A connection pool (HikariCP for the database) lends a fixed number of connections; when all are lent out, the next thread waits - up to connection-timeout, 30 s by default - and then usually gets one. So exhaustion shows as latency, not errors: the only error is Connection is not available, request timed out after 30000ms, and only after 30 s.
Confirm: hikaricp_connections_active = max and hikaricp_connections_pending > 0; a thread dump full of threads in HikariPool.getConnection.
The usual cause is holding a connection while waiting on something else - a @Transactional method that calls another service over HTTP. Little's law: 20 connections held 8 s each = 2.5 requests/s for the whole pool, for every endpoint that needs the database.
The fix is a timeout on the slow call and not holding the connection across it - not a bigger pool: the database has finite parallelism, and 10 pods x 200 connections can exceed max_connections. leak-detection-threshold shows who holds connections too long.
Also asked: What is a connection pool and why do applications use one? · Why is a bigger connection pool often slower? · How would you size a database connection pool for a service with several replicas?
Learn it: 21.15 Connection pools: why exhaustion is latency, not errors
What timeouts should an HTTP client have, and how do you choose them? Mid
- Connect - establishing TCP (and TLS): "is anyone there?". Fails fast when the host is down or a firewall drops.
- Read - the longest silence once connected: "is it still answering?". This is the one that protects your threads - without it a call to a hung service waits forever.
- Total - the whole operation including retries.
- Pool wait - waiting for a pooled connection before any of the above.
Assume there is no timeout unless someone set it: HttpURLConnection and Reactor Netty have no read/response timeout by default, and on oncall-lab orders had http.client.read.timeout.ms=0 (infinite).
Choosing values - the budget: set each from the dependency's latency (a few times its p99), make the sum on the sequential path (times retries) fit inside your own SLO, propagate the deadline (pass on only what is left), and keep upstream timeouts longer than downstream ones - nginx proxy_read_timeout > the app's total > each client's read timeout.
Also asked: Explain timeout budgets across a chain of service calls. · A dependency hung and your service went down with it. Explain the mechanism. · What does a connect timeout tell you that a read timeout does not?
Learn it: 21.18 Timeouts: connect, read, total - and the budget
What is a circuit breaker and how does it work? Mid
A circuit breaker stops calling a dependency that is failing, so you fail fast (and free the thread) instead of waiting on it:
- CLOSED - calls go through; outcomes are recorded in a sliding window.
- Failure rate (or slow-call rate) over the threshold, after a minimum number of calls -> OPEN: every call is rejected immediately (
CallNotPermittedException), and you answer with a fallback (a cached value, "payment pending") or a fast 503. - After
wait-duration-in-open-state-> HALF_OPEN: a few trial calls; success closes it, failure opens it again.
In Resilience4j it is properties per instance (sliding-window-size, minimum-number-of-calls, failure-rate-threshold). The defaults (100 calls, 60 s slow-call threshold) are slow to trip on a low-traffic service - tune them.
The catch: it needs a timeout underneath. It counts a call when the call completes; a hung call with no read timeout never completes, so the breaker never opens. Timeout first, then breaker, plus a bulkhead so only a bounded number of threads can wait on that dependency.
Also asked: When is it safe to retry a failed request? · Explain retry storms and how jitter prevents them. · What is the bulkhead pattern?
Learn it: 21.22 Retries, jitter, circuit breakers, bulkheads, backpressure
An async service has high latency, no errors and idle Tomcat threads. How do you investigate? Mid
With a CompletableFuture controller the waiting did not go away - it moved from Tomcat's threads to an executor. So I look at the executors:
- Thread dump: the executor's threads (
orders-async-1..8) allRUNNABLEinNet.pollinside the same client call (ShippingClient.eta), whilehttp-nio-8080-exec-*wait inTaskQueue.take- Tomcat looks idle. - Metrics:
executor_active_threadsat the maximum andexecutor_queued_tasksgrowing. Nothing throws: tasks queue in an unboundedLinkedBlockingQueueuntil the async request timeout answers 503 - and the task still runs later for a client that has gone.
Then check the code for supplyAsync without an executor: it uses ForkJoinPool.commonPool() (CPUs - 1, shared by the whole JVM), and on 2 CPUs or fewer a new thread per task. Nobody chose the concurrency.
Fixes: a named, bounded executor per dependency, a bounded queue (reject fast), a timeout inside the task (orTimeout or a read timeout), and alerts on the queue metric.
Also asked: What is the risk of CompletableFuture.supplyAsync without an executor argument? · Why is an unbounded queue in front of a thread pool dangerous? · What is the difference between synchronous and asynchronous request handling in a web service?
Learn it: 21.25 Async and CompletableFuture: when nothing throws
What is a Python virtual environment, and how would you install a small Python tool on a server? Mid
A venv is a directory with its own bin/python (a link to the system interpreter) and its own site-packages - packages for one project only, like a project-local node_modules. Nothing global changes.
Why it matters on a server: system Python belongs to apt. Tools like unattended-upgrades and cloud-init depend on exact versions of its packages, so sudo pip install can break the OS weeks later. Ubuntu blocks it (PEP 668, "externally managed environment"); --break-system-packages is named honestly.
Installing a tool:
python3 -m venv /opt/tool/.venv(needspython3-venv)./opt/tool/.venv/bin/pip install -r requirements.txt- a list pinned withpip freeze(exact versions = a lockfile).- Run it by the venv's interpreter path:
ExecStart=/opt/tool/.venv/bin/python /opt/tool/check.py. systemd and cron do notactivateanything;activateonly prepends.venv/binto PATH.
Never commit the venv; recreate it from the pinned list. For CLI tools you just want to run: pipx.
Also asked: Why is running sudo pip install on a production server dangerous? · How do you pin a Python tool's dependencies? · How would you run a Python script from a systemd timer so it uses the right packages?
Learn it: 21.29 Python for operations: the subset, venv, and why never system pip
What makes an operations script production-ready? Mid
Each piece prevents one failure of an unattended script (cron, a systemd timer, a CI job):
- argparse - real options and
--help; usage errors exit 2. - logging, not print - timestamps, levels, a logger name,
-vfor detail; to stdout/stderr so journald or the container runtime collects it. - Timeouts on everything -
requests.get(url, timeout=(1, 5)); requests has no default timeout, so a health checker without one hangs on exactly the service that is down.subprocess.run(..., timeout=5)too. - Commands as a list, never
shell=Truebuilt from input (shell injection), withcheck=Trueso a non-zero exit raises instead of passing silently - Python'sset -e. - Meaningful exit codes -
sys.exit(main()): 0 ok, 1 unexpected, 2 usage, 3+ documented. Expected failures as a custom exception caught at the top; real bugs exit 1 with a traceback. - Retries only where safe - a
Sessionwith aRetryadapter, idempotent methods only.
Also asked: Why is subprocess with shell=True dangerous? · What happens when you call requests.get without a timeout? · How should a script report failure to cron or systemd?
Learn it: 21.31 An ops script done properly: argparse, logging, requests, subprocess, exit codes
How do you test a Python ops script that calls the network or other systems? Mid
With pytest: files test_*.py, functions test_*, plain assert (pytest rewrites it to show the values on failure). Run pytest -v, -x to stop at the first failure, -k name to select.
Keep the logic in plain functions (status_to_code, main(argv) returning an exit code) so it can be called from a test, and replace the edges with fixtures:
monkeypatch.setattr(health, "fetch_status", lambda url, timeout: "DOWN")- no network needed; undone after the test.tmp_path- a fresh directory for config and output files.capsys- capture what was printed.
Then assert health.main(["--url", "http://x"]) == 2 tests the exit code a cron job or a pipeline would see. pytest's own exit code (0 passed, 1 failed, 5 no tests) is what a build step relies on.
For Kubernetes: the Python client with config.load_kube_config() locally and load_incluster_config() in a pod, with a ServiceAccount Role that allows only what the tool needs.
Also asked: How should a tool running inside a pod authenticate to the Kubernetes API? · What is a pytest fixture, and which built-in ones do you use most? · How would you structure a Python ops tool so that it is easy to test?
Learn it: 21.34 pytest, fixtures, and the Kubernetes Python client
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.