Why this chapter exists
It is 3am. A shop's checkout has been failing for a third of its customers for twenty minutes - an outage (a period when a service does not work for its users, fully or partly). Nobody was woken up, because the only alarm the team had watched the wrong number. When someone finally looks, they cannot say how bad it is, whether it is getting better, or when it started.
SRE (Site Reliability Engineering: the job of keeping a software service working for its users, with measurement and automation instead of heroics) is the answer to that night. This chapter gives you its vocabulary - the handful of numbers and habits every SRE team uses - and makes you compute each one yourself from the records a real service leaves behind, after a real incident (an unplanned event that hurts users; each one gets a number, like #4471).
What you need to know already: nothing. This is the first chapter. Every word and every command is explained the first time it appears.
A few words first
A service here means a program that other programs or people send work to - say, "checkout", which takes a shopping cart and turns it into an order. It runs on a server (a computer, usually in a data centre, that runs services for others and has no screen or keyboard of its own).
A user of the service sends it a request (one message asking for one thing: "show my cart", "pay for this order") and gets back a response. On the web, requests and responses use HTTP (the protocol, or agreed message format, that browsers and web services speak). Every HTTP response carries a three-digit status code that says how it went:
2xx success 200 OK, 201 Created
3xx go elsewhere 301 Moved
4xx the CLIENT got it wrong 404 Not Found, 401 not logged in
5xx the SERVER failed 500 Internal Server Error, 502/504 (a proxy gave up)
"5xx" means "any code from 500 to 599". For reliability, 5xx is what counts: the service failed the user.
In front of checkout sits nginx (a popular web server that receives every request first and passes it on to the right service - a proxy). nginx writes one line per request to an access log (a text file where a program records what it did, one event per line). That log is the raw material of this chapter.
If you can only measure four things
Google's SRE book (chapter 6, Monitoring Distributed Systems) says: if you can only measure four metrics (numbers measured over and over, so you can watch them change) of a user-facing service, measure these - the four golden signals.
latency how long requests take
traffic how much demand is arriving
errors how many requests fail
saturation how full the service is
For a web service, concretely:
| Signal | Concrete metric |
|---|---|
| Latency | p50 / p95 / p99 of request duration (explained below), successes and failures kept apart |
| Traffic | requests per second, per route (the address a request asks for, like /api/cart) if routes differ a lot |
| Errors | fraction of requests returning 5xx (the error rate), plus "wrong" answers (a 200 whose body says "error") and "too slow" answers if you promised a speed |
| Saturation | the most constrained resource (anything the service needs a share of to work): how many of the service's workers are busy, how many database connections are in use, memory, open files, queue length |
Latency: percentiles, never averages
Latency is the time between a request arriving and its response leaving.
An average is dominated by the few slowest requests and describes nobody. A service that answers 99 requests in 50 ms (milliseconds: thousandths of a second) and one in 60 s has a mean of about 650 ms - a number no request actually experienced. Averages also hide bimodal distributions (two separate groups of values): fast answers from a cache plus slow answers from the database average out to a "normal" number nobody got.
A percentile is a position in the sorted list of values. Sort every request's duration from fastest to slowest; the p99 (99th percentile) is the value that 99% of requests are at or below.
p50 the median user: half the requests are faster, half slower
p95 one request in twenty is slower than this
p99 one request in a hundred is slower than this
The slow end of the list - the tail, or tail latency - matters more than it looks. If loading one screen of an app needs 100 calls to other services (a backend is a service that another service calls behind the scenes), and each has a 1% chance of being slow, the chance that the screen is slow is 1 - 0.99^100 = 63%. Your p99 is your users' median.
Two more rules:
- Separate the latency of errors from the latency of successes. A fast 500 drags the average down and makes things look better. A slow 500 (a timeout: the caller stopped waiting after a fixed time) is the worst of both.
- Percentiles do not average. You cannot average the p99 of five copies of a service into one overall p99; you need the underlying values. The next lesson shows why.
Saturation: the one people forget
Latency, traffic and errors all describe what already happened. Saturation is the leading indicator (a number that moves before the trouble shows): how close you are to the point where latency and errors fall off a cliff. It is forgotten because it is not in the access log. You have to go and look at the resource that runs out first.
Most systems degrade before 100%: a queue that is 80% full already adds wait time, and a connection pool (a fixed set of open connections to a database that requests borrow and hand back) at its maximum makes every extra request wait for a free slot. That is why you alert on saturation with a threshold below 100%, or on a prediction ("the disk fills in 4 hours").
For a typical web service, the resources that run out are rarely the processor:
- workers - the service can handle, say, 200 requests at once; the 201st waits
- the database connection pool - how many connections are in use out of the maximum, and how many requests are pending (waiting for one)
- memory - how full it is, and whether the program spends its time cleaning up memory instead of serving requests
- open files - every file and every network connection a program holds open takes one slot out of a fixed limit (often 1024)
A service can be at 15% CPU (the CPU, central processing unit, is the chip that executes the program; 15% busy is nearly idle) and completely saturated: every worker is waiting for a database connection that is not coming.
Why these four and not more
They answer the only questions that matter when you are woken up: is it broken for users (latency, errors), how much is being asked of it (traffic), and is it about to break (saturation). Everything else is a cause you look at during diagnosis, not a signal you watch all day. More top-level signals means more dashboards (screens of graphs) nobody reads and more alerts (automatic messages that fire when a number crosses a line) nobody trusts.
Two related acronyms you will hear:
RED Rate, Errors, Duration per service (golden signals minus saturation)
USE Utilisation, Saturation, Errors per resource (CPU, disk, pool, network card)
RED tells you a service hurts; USE tells you which resource is the reason.
The terminal, in five minutes
Every number in this chapter comes from files on this box (the simulated server on the left; it runs Linux, an operating system - the base software that runs a computer and every program on it). You talk to it through the terminal: a window where you type text commands and read their text output. The program reading what you type is the shell (here bash); it shows a prompt like learner@oncall-lab:~$ when it is ready for the next command.
A command is a program name followed by arguments (the things you give it, separated by spaces). Arguments that start with - are options, also called flags: they switch a behaviour on. head -1 access.log runs the program head with the flag -1 ("only one line") and the argument access.log (which file).
Files live in directories (folders). A path says where a file is: /var/log/nginx/access.log starts at the top of the whole disk (/), goes into var, then log, then nginx. ~ is short for your own home directory.
The commands this chapter leans on, in one line each:
cat FILE print a whole file
head -1 FILE print its first line (head -3: the first three)
tail -1 FILE print its last line
wc -l count lines ("word count", -l = lines only)
grep TEXT FILE print only the lines that contain TEXT
sort -n sort lines as numbers (-n); without -n it sorts as text
uniq -c collapse repeated neighbouring lines, -c = with a count
awk '...' a small language for columns: $1 is the first column,
$NF the last, NR the line number; {s+=$1} END {print s} sums
jq '...' the same idea for JSON (a text format of {"key": value} records)
Two pieces of glue:
A | Bis a pipe: the output of command A becomes the input of command B.grep 504 access.log | wc -lcounts the lines containing 504.A > fileis a redirection: the output goes into a file instead of the screen (and replaces what was in it).
That is enough to start. Each mission says what the commands it needs are for; the hints get more concrete, and solution shows a full worked session. You will learn all of these properly in the Linux chapters.
Later (Ch 7): grep, sort, uniq, awk and jq get a whole chapter; here you only need the one-liners shown.
Doing it with the tools you have
This box's nginx writes the request duration ($request_time, in seconds) as the last column of every access-log line, and the status code as column 9. So the access log already contains three of the four signals:
# traffic: count the lines (one line = one request) and look at the first and last time
head -1 access.log; tail -1 access.log; wc -l < access.log
# errors: share of 5xx. For each line whose column 9 is 500 or more, add one to e;
# at the END print e as a percentage of NR (the number of lines read)
awk '$9 >= 500 {e++} END {printf "%.1f%%\n", 100*e/NR}' access.log
# latency: print every duration in ms, sort them as numbers, then pick the one at rank 99%
awk '{print $NF*1000}' access.log | sort -n |
awk '{a[NR]=$1} END {i=int(NR*0.99); if (i<NR*0.99) i++; print "p99", a[i]}'
(; separates commands on one line; < access.log feeds the file to wc as input so it prints only the number; printf prints with a format, %.1f = one decimal, %% = a literal percent sign.)
The last pipeline sorts the values and picks the value at position 99% of the way down the list, rounded up: that is the nearest-rank percentile. Monitoring tools estimate percentiles slightly differently, so their numbers differ a little; the idea is the same.
What you can now do:
- name the four golden signals and a concrete number for each
- explain why an average latency describes nobody, and what p50 and p99 mean
- read a simple command line: program, flags, arguments, pipe