SRE Fundamentals: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 0 of the course.
What are SLIs, SLOs and error budgets, and how do they relate? Junior
- SLI (service level indicator): a measurement of what users get, written as good events / valid events. For checkout: the share of valid requests (health checks left out) that did not return a 5xx.
- SLO (objective): a target for that SLI over a window, for example 99.9% over 30 days. Never 100%: users cannot tell, it costs a fortune and it forbids any change.
- Error budget: 1 - SLO, the failure you are allowed. 99.9% over 30 days is 0.1% of requests, or 43.2 minutes of full outage.
- How they connect: incidents and risky deploys spend the budget. With budget left, the team ships; when it is gone, the error budget policy agreed in advance says reliability work comes first. An SLA is the looser contract with customers, with money attached if it is missed.
Also asked: Why is 100% the wrong reliability target? · What is the difference between an SLO and an SLA? · What would you put on the first dashboard for a web service?
What are the four golden signals? Junior
They are the four things to measure first on a user-facing service, from the Google SRE book: latency, traffic, errors and saturation.
- Latency: how long requests take, as percentiles (p50, p95, p99), never an average, with successes and failures kept apart.
- Traffic: how much demand arrives, usually requests per second.
- Errors: the share of requests that fail - 5xx, plus wrong answers (a 200 whose body says "error").
- Saturation: how full the most constrained resource is - busy workers, database connections in use, memory, open files.
The first three describe what already happened. Saturation is the leading indicator: it moves before latency and errors fall off a cliff. From an nginx access log you get the first three with wc -l, awk '$9 >= 500' and sort -n; saturation needs the service's own numbers.
Also asked: Why is the average latency a bad number to put on a dashboard? · What does p99 mean, and why look at it as well as p50? · What is the difference between RED and USE?
Learn it: 0.1 The four golden signals
How do you calculate a percentile, and why is the mean not enough? Junior
Sort all the values as numbers and pick the one at a rank. With nearest rank, the rank is ceil(N x q), counting from 1: for ten requests of 9 10 11 12 13 14 15 16 30 900 ms, p50 is the 5th value (13 ms) and p90 the 9th (30 ms).
The mean of those ten is 103 ms: slower than nine of them and far faster than the tenth, so it describes nobody. A few slow requests drag it up, and it hides two groups of values (fast and slow) behind one "normal" number.
On a box the recipe is sort -n | awk: print one number per request, sort numerically, pick the rank, rounding up. Two classic mistakes: plain sort sorts as text (1000 before 120), and averaging p99s from several instances or hours - percentiles do not average, you need the underlying values.
Also asked: Why can't you average the p99 of five instances to get the overall p99? · What is a histogram, and how do you estimate a percentile from one? · Why should error latency be kept apart from success latency?
Learn it: 0.2 Percentiles by hand, and why they do not average
What is the USE method? Junior
A checklist for finding which resource is the problem: for every resource, check Utilisation, Saturation and Errors.
- Utilisation: how busy it is - percent busy, or used out of capacity.
- Saturation: how much work is waiting for it - a queue.
- Errors: failures - "too many open files", timeouts, error counters.
On a box I walk it resource by resource: CPU with vmstat 1 2 (r = waiting for a CPU), memory with free -m (the available column, and swap in use), disk with df -h, open files against their limit, and then the application's own resources, like a database pool's connections in use out of its maximum. The golden signals say a service hurts; USE says which resource is the reason. In checkout's outage the pool was at 30/30 with requests waiting while the CPU was nearly idle.
Also asked: What is the difference between utilisation and saturation? · A service is slow but its CPU is at 15%. What else could be saturated? · What does Little's law tell you about a connection pool?
Learn it: 0.5 Saturation and the USE method
What is the difference between an SLI, an SLO and an SLA? Junior
- SLI, a measurement: good events / valid events. For a web API, the share of valid requests (health checks left out) that did not return a 5xx, or the share answered within 300 ms.
- SLO, a target for an SLI over a window: 99.9% over 30 days. It is the team's own goal, and it defines the error budget (1 - SLO; 43.2 minutes a month at 99.9%).
- SLA, a contract with customers with a consequence, usually money back, if it is missed. It is set looser than the SLO, so the team reacts long before the contract is at risk.
A quick test: if missing it costs money, it is an SLA; if it changes what the team works on, it is an SLO. And never aim for 100%: users cannot tell 99.99% from 100%, and it would forbid every change.
Also asked: Give an example of an availability SLI and a latency SLI for a web API. · How much downtime does each extra nine allow over 30 days? · Why is the SLA looser than the SLO?
Learn it: 0.8 SLIs, SLOs and SLAs
Should health checks be included in an availability SLI? Junior
No. An SLI counts valid events - users trying to do something - and health checks are not users. They hit a cheap address like /healthz that almost never fails, so counting them makes the service look better than it is: on checkout's log, adding the probe raised availability from 96.00% to 96.67%. They also pad the denominator, so a real outage looks smaller.
So the SLI filters them out, for example with jq 'select(.uri != "/healthz")', and the filter is written into the SLI's definition so people can review it. The same goes for other traffic that is not a user. When you combine days or routes, use the ratio of sums (all good / all valid), never the mean of the daily percentages.
Also asked: Where would you measure an SLI: at the proxy in front, or inside the service, and why? · What is the difference between a rolling and a calendar window? · Why is "ratio of sums" right and "mean of ratios" wrong when combining SLIs?
Learn it: 0.9 Choosing what to count: valid events, measurement points, windows
What is an error budget and how is it used? Junior
The error budget is 1 - SLO: how much failure the service is allowed over the window. For 99.9% over 30 days that is 0.1% of valid requests, or about 43 minutes of full outage.
Everything risky spends it: incidents, bad releases, maintenance, experiments. While budget is left, the team ships; the reliability users need is already paid for. When it runs out, an error budget policy agreed in advance applies - typically feature releases pause and the team works on reliability until the service is back within its SLO. That turns the old "dev wants to ship, ops wants stability" argument into arithmetic.
How fast it is spent is the burn rate: error ratio / budget fraction. A 1% error rate against a 99.9% SLO is a burn rate of 10 - the monthly budget gone in three days.
Also asked: How much downtime does a 99.9% SLO allow per month? · What is a burn rate, and what does a burn rate of 1 mean? · Who decides what happens when the error budget runs out?
Learn it: 0.16 Error budgets
How much downtime does a 99.9% SLO allow per month? Junior
About 43 minutes over 30 days: 30 x 24 x 60 = 43,200 minutes, and 0.1% of that is 43.2. Scale from there: 99% is ten times more (7.2 hours), 99.99% ten times less (about 4.3 minutes).
Two things I'd add. Most incidents are partial, so convert with duration x error ratio: 37 minutes at 30% errors uses about 11 of the 43 minutes. And the budget is usually counted in requests, not minutes: 0.1% of the month's valid requests, so an hour of errors at peak traffic costs more than an hour at 4am, which matches what users felt. To see how long is left, divide the budget remaining by the current burn rate - time left, not the 30 days total.
Also asked: A 20-minute incident failed 25% of requests. How much of a 99.9% monthly budget did it use? · How do you find which days or events used up the error budget? · Why does a rolling 30-day window take so long to recover after a big incident?
Learn it: 0.17 Budget arithmetic: minutes, requests, burn rate, time left
Why should you alert on symptoms rather than causes? Junior
Because pages should fire when users are hurt, and only then.
Most causes do not hurt anyone: CPU at 85% during a batch job, a restart nobody noticed, a disk at 71% with weeks of room. Paging on them wakes people for nothing, and people learn to ignore pages (alert fatigue). And you cannot list every cause in advance, but almost all of them show up as the same symptom: users getting errors or slow answers. Checkout's outage happened at 18% CPU, so a CPU alert would have stayed silent.
So the page is a symptom, ideally the SLO's burn rate: "checkout is burning its 99.9% budget more than 14.4 times too fast over the last hour and the last 5 minutes". Causes go on the dashboard the page links to, and slow problems become tickets, like "disk full in four days".
Also asked: What makes an alert actionable? · When should something be a page, a ticket, or only a dashboard? · Why does a burn-rate alert check two windows?
Learn it: 0.21 Alerting philosophy, and multi-window burn rates
A burn-rate alert keeps firing for an hour after the outage was fixed. Why, and how do you fix it? Junior
The rule probably looks only at a long window. A 1-hour error ratio still contains the outage's minutes for up to an hour after the fix, so the burn rate stays above the threshold while users are already fine. People get used to stale pages and start ignoring them.
The fix is the multi-window rule: pair the long window with a short one, 1/12 of its length (5 minutes for 1 hour, 30 minutes for 6 hours), and fire only when both are above the burn rate (14.4 for the fast page). After the fix the 5-minute ratio drops to zero within five minutes, the and turns false and the alert resolves. The long window still makes sure it only fires for real budget loss. I'd also check for a long for: on the rule, which delays both firing and clearing.
Also asked: How long does a burn-rate rule take to fire for a 100% outage, and for a 2% error rate? · How do you alert for a service that only gets a few requests an hour? · How would you measure whether an alert is any good?
Learn it: 0.22 Designing burn-rate alerts: detection, reset, low traffic, noise
What roles are there in incident response, and why? Junior
- Incident commander (IC): owns the incident. Keeps the big picture, assigns roles, makes the calls (roll back or not, call in more people). Does not debug, because someone has to watch the whole.
- Ops lead: hands on keyboard, the only one changing production, so two people's fixes do not collide.
- Comms lead: the status page and the stakeholders, updates on a fixed schedule, and keeps "any update?" messages away from the ops lead.
- Scribe (or the incident channel itself): the timeline - every decision, action and observation, with UTC times. It becomes the postmortem.
Without roles everyone debugs, nobody talks to customers and nobody can say later what happened when. In a small team one person holds several roles; the first split is IC from ops lead, and handovers are said out loud.
Also asked: What does a SEV-1 mean compared with a SEV-3? · Why mitigate first and look for the cause later? · What goes into a status update during an incident?
Learn it: 0.29 Running an incident: severity, roles, comms, and the clock
What is a blameless postmortem? Junior
A written review after an incident that asks how the system let the failure happen, not who failed. It assumes everyone acted reasonably with what they knew at the time.
Concretely: it names roles ("the on-call"), not people; it describes what people saw and why their action made sense then; it avoids hindsight words like "should have" and "forgot"; and "human error" is never the cause - if a person could make the mistake, a guardrail was missing. Its sections: summary, impact in numbers (including error budget), a UTC timeline, contributing factors, detection, resolution, action items with owner and ticket, and lessons learned.
Why: the facts live in people's heads. If honesty gets people blamed, the accounts get short and defensive, near misses stop being reported, and the real conditions stay in place for the next person.
Also asked: What makes a postmortem action item good, and what makes one useless? · What is toil? Give an example. · Why does a postmortem include "where we got lucky"?
Learn it: 0.30 Blameless postmortems, toil, and running an incident
How do you build an incident timeline? Junior
From evidence first, memory last.
I collect the sources: the access log (first and last failed request = when users were hurt), the error logs, the journal (deploys, sudo commands), the pager export (fired, acknowledged, resolved) and the chat export (what people knew and decided). Each uses its own time format, so I convert every line to the same ISO form in UTC, tag it with its source, and merge them with sort - ISO timestamps sort correctly as text.
Rules: always UTC; impact starts at the first failed request, not when someone noticed; one event per line; include decisions and observations, not only actions. Then I read the gaps: in #4471 there were 11 minutes between declaring SEV-2 and the first rollback attempt. Explaining the gaps is where the findings are.
Also asked: Rewrite "the engineer forgot to update the runbook" in blameless language. · What is the difference between a trigger and a contributing factor? · Why is "human error" not accepted as a cause?
Learn it: 0.31 Writing the postmortem: evidence, timeline, blameless language
What is toil, and how is it different from other operational work? Junior
Toil, as the SRE book defines it, is work that is manual, repetitive, automatable, tactical (interrupt-driven), has no enduring value - the service is no better afterwards - and grows with the service. Restarting the same crashing worker every night is the classic example.
It is not overhead (meetings, planning, training), which is necessary but not operational, and it is not engineering (automation, SLOs, postmortems, fixing a root cause, even a painful one-off migration). Unpleasant is not the same as toil: if the service is permanently better afterwards, it was engineering.
Why it matters: toil scales with the service and engineering does not, so Google caps SRE toil at 50% of the time. To choose what to remove first, I measure it from the ticket export, rank tasks by hours per year, and prefer eliminating a task over automating it.
Also asked: How do you decide whether a task is worth automating? · Why is "eliminate" better than "automate"? · What would you check before trusting a cleanup script that runs on a schedule?
Learn it: 0.35 Measuring toil, and choosing what to remove first
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.