OnCallReady

Chapter 0 SRE Fundamentals

Golden signals, SLIs and SLOs, error budgets, burn-rate alerting, blameless postmortems and toil - computed from real logs, not recited.

In plain words

Think of a bus company. Passengers don't care how shiny the engines are; they care whether the bus comes, whether it arrives on time, and whether it breaks down. So the company promises "95 of every 100 buses on time", measures it honestly, and treats the 5 late buses a month as an allowance: spend it on road works and new routes, but if you use it up by the 10th, stop experimenting and fix the buses. When a bus does break down, one person coordinates, nobody gets blamed, and afterwards everyone writes down what to change so it doesn't happen again.

Site Reliability Engineering is that way of running software. The four golden signals and USE are how you measure, SLIs, SLOs and error budgets are the promise and the allowance, burn-rate alerts wake you only when the allowance is running out fast, and incidents, postmortems and toil reduction are how the team gets better.

Why it matters on call

This chapter is the vocabulary of the job you are moving into. In an SRE or platform interview you will be asked to define an SLO for an API, explain why you alert on burn rate instead of CPU, compute how much budget a 37-minute partial outage used, and describe a blameless postmortem. On the job, you will be paged at 3am by a burn-rate alert, run the incident clock, and argue with data that a feature freeze is the policy working rather than someone's opinion.

It comes first, before any Linux, because it tells you what all the later tools are for: every command you learn afterwards is a way to measure, alert on or fix one of these things. You still compute every number yourself from a real box - small text tools (awk, jq) over web server logs, the journal for the timeline, pager exports for detection time - and each mission explains the commands it needs.

Later (Ch 27): the monitoring tools that compute these numbers for you continuously.

Lessons

  1. The four golden signals
  2. Percentiles by hand, and why they do not average
  3. Saturation and the USE method
  4. SLIs, SLOs and SLAs
  5. Choosing what to count: valid events, measurement points, windows
  6. Error budgets
  7. Budget arithmetic: minutes, requests, burn rate, time left
  8. Alerting philosophy, and multi-window burn rates
  9. Designing burn-rate alerts: detection, reset, low traffic, noise
  10. Running an incident: severity, roles, comms, and the clock
  11. Blameless postmortems, toil, and running an incident
  12. Writing the postmortem: evidence, timeline, blameless language
  13. Measuring toil, and choosing what to remove first

25 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

Is SRE just a new name for operations?

No. Classic operations runs systems by hand and grows with them. SRE, as Google defined it, applies software engineering to operations: reliability is a measured product feature (SLOs), the acceptable failure rate is an explicit budget shared with product, toil is capped at around half the team's time and actively automated away, and incidents are learned from blamelessly. Many companies say "SRE" for platform or DevOps roles, but these ideas are what interviewers probe.

What is the difference between SLI, SLO and SLA in one line each?

The SLI is the measurement: good events divided by valid events, like "96.0% of checkout requests succeeded". The SLO is the internal target for it over a window: "99.9% over a rolling 30 days". The SLA is a contract with a customer and a consequence, usually looser than the SLO: "below 99.5% in a month, 10% credit". Engineers own SLOs; legal and sales own SLAs.

Why alert on SLO burn rate instead of CPU or memory?

Because users feel symptoms, not causes. CPU at 85% during a batch job hurts nobody, and an outage from a shrunken connection pool can happen at 18% CPU. A burn-rate alert fires when user-facing failures are eating the error budget fast enough to matter, which catches causes you never predicted and ignores those that don't hurt. Causes stay on dashboards for diagnosis and become tickets when they predict trouble.

Why does this chapter compute everything by hand with awk and jq instead of a monitoring tool?

So you understand what the tools compute. A p99 is a sort and a rank, an SLI is a ratio of sums, a burn rate is an error ratio divided by the budget. Once you have done it by hand on web server logs, any monitoring system is just automation of arithmetic you already trust, and you can tell when a dashboard is lying. The terminal commands are explained in each mission; you do not need to know them yet.

Is blameless the same as no accountability?

No. Blameless means people are not treated as the cause of a failure, because the system allowed the action, and nobody is punished for what a postmortem reveals. Accountability moves to the future: action items have named owners, priorities, tickets and done-when criteria, and they are tracked until they ship. Blame destroys the honest information you need; ownership of fixes does not.