OnCallReady

Lesson 0.1 · SRE Fundamentals · 19 min read

The four golden signals

In plain words

Imagine checking on a pizza shop by standing outside for a minute. You notice four things: how long people wait for their pizza (latency), how many people walk in (traffic), how many walk out angry with the wrong order (errors), and whether the kitchen looks overloaded, with orders piled up by the oven (saturation). With just those four, you know if the shop is fine, busy, broken, or about to be.

The four golden signals, from the Google SRE book, are the same for a service: latency as percentiles (p50, p95, p99) kept separate for successes and failures, traffic in requests per second, errors as the fraction of 5xx (plus wrong answers), and saturation as the most constrained resource, like the pool of database connections or the worker threads. The web server's access log (one line per request, including how long it took) already holds three of them.

Why this chapter exists

It is 3am. A shop's checkout has been failing for a third of its customers for twenty minutes - an outage (a period when a service does not work for its users, fully or partly). Nobody was woken up, because the only alarm the team had watched the wrong number. When someone finally looks, they cannot say how bad it is, whether it is getting better, or when it started.

SRE (Site Reliability Engineering: the job of keeping a software service working for its users, with measurement and automation instead of heroics) is the answer to that night. This chapter gives you its vocabulary - the handful of numbers and habits every SRE team uses - and makes you compute each one yourself from the records a real service leaves behind, after a real incident (an unplanned event that hurts users; each one gets a number, like #4471).

What you need to know already: nothing. This is the first chapter. Every word and every command is explained the first time it appears.

A few words first

A service here means a program that other programs or people send work to - say, "checkout", which takes a shopping cart and turns it into an order. It runs on a server (a computer, usually in a data centre, that runs services for others and has no screen or keyboard of its own).

A user of the service sends it a request (one message asking for one thing: "show my cart", "pay for this order") and gets back a response. On the web, requests and responses use HTTP (the protocol, or agreed message format, that browsers and web services speak). Every HTTP response carries a three-digit status code that says how it went:

2xx   success          200 OK, 201 Created
3xx   go elsewhere     301 Moved
4xx   the CLIENT got it wrong    404 Not Found, 401 not logged in
5xx   the SERVER failed          500 Internal Server Error, 502/504 (a proxy gave up)

"5xx" means "any code from 500 to 599". For reliability, 5xx is what counts: the service failed the user.

In front of checkout sits nginx (a popular web server that receives every request first and passes it on to the right service - a proxy). nginx writes one line per request to an access log (a text file where a program records what it did, one event per line). That log is the raw material of this chapter.

If you can only measure four things

Google's SRE book (chapter 6, Monitoring Distributed Systems) says: if you can only measure four metrics (numbers measured over and over, so you can watch them change) of a user-facing service, measure these - the four golden signals.

latency      how long requests take
traffic      how much demand is arriving
errors       how many requests fail
saturation   how full the service is

For a web service, concretely:

SignalConcrete metric
Latencyp50 / p95 / p99 of request duration (explained below), successes and failures kept apart
Trafficrequests per second, per route (the address a request asks for, like /api/cart) if routes differ a lot
Errorsfraction of requests returning 5xx (the error rate), plus "wrong" answers (a 200 whose body says "error") and "too slow" answers if you promised a speed
Saturationthe most constrained resource (anything the service needs a share of to work): how many of the service's workers are busy, how many database connections are in use, memory, open files, queue length

Latency: percentiles, never averages

Latency is the time between a request arriving and its response leaving.

An average is dominated by the few slowest requests and describes nobody. A service that answers 99 requests in 50 ms (milliseconds: thousandths of a second) and one in 60 s has a mean of about 650 ms - a number no request actually experienced. Averages also hide bimodal distributions (two separate groups of values): fast answers from a cache plus slow answers from the database average out to a "normal" number nobody got.

A percentile is a position in the sorted list of values. Sort every request's duration from fastest to slowest; the p99 (99th percentile) is the value that 99% of requests are at or below.

p50   the median user: half the requests are faster, half slower
p95   one request in twenty is slower than this
p99   one request in a hundred is slower than this

The slow end of the list - the tail, or tail latency - matters more than it looks. If loading one screen of an app needs 100 calls to other services (a backend is a service that another service calls behind the scenes), and each has a 1% chance of being slow, the chance that the screen is slow is 1 - 0.99^100 = 63%. Your p99 is your users' median.

Two more rules:

Saturation: the one people forget

Latency, traffic and errors all describe what already happened. Saturation is the leading indicator (a number that moves before the trouble shows): how close you are to the point where latency and errors fall off a cliff. It is forgotten because it is not in the access log. You have to go and look at the resource that runs out first.

Most systems degrade before 100%: a queue that is 80% full already adds wait time, and a connection pool (a fixed set of open connections to a database that requests borrow and hand back) at its maximum makes every extra request wait for a free slot. That is why you alert on saturation with a threshold below 100%, or on a prediction ("the disk fills in 4 hours").

For a typical web service, the resources that run out are rarely the processor:

A service can be at 15% CPU (the CPU, central processing unit, is the chip that executes the program; 15% busy is nearly idle) and completely saturated: every worker is waiting for a database connection that is not coming.

Why these four and not more

They answer the only questions that matter when you are woken up: is it broken for users (latency, errors), how much is being asked of it (traffic), and is it about to break (saturation). Everything else is a cause you look at during diagnosis, not a signal you watch all day. More top-level signals means more dashboards (screens of graphs) nobody reads and more alerts (automatic messages that fire when a number crosses a line) nobody trusts.

Two related acronyms you will hear:

RED   Rate, Errors, Duration           per service  (golden signals minus saturation)
USE   Utilisation, Saturation, Errors  per resource (CPU, disk, pool, network card)

RED tells you a service hurts; USE tells you which resource is the reason.

The terminal, in five minutes

Every number in this chapter comes from files on this box (the simulated server on the left; it runs Linux, an operating system - the base software that runs a computer and every program on it). You talk to it through the terminal: a window where you type text commands and read their text output. The program reading what you type is the shell (here bash); it shows a prompt like learner@oncall-lab:~$ when it is ready for the next command.

A command is a program name followed by arguments (the things you give it, separated by spaces). Arguments that start with - are options, also called flags: they switch a behaviour on. head -1 access.log runs the program head with the flag -1 ("only one line") and the argument access.log (which file).

Files live in directories (folders). A path says where a file is: /var/log/nginx/access.log starts at the top of the whole disk (/), goes into var, then log, then nginx. ~ is short for your own home directory.

The commands this chapter leans on, in one line each:

cat FILE          print a whole file
head -1 FILE      print its first line   (head -3: the first three)
tail -1 FILE      print its last line
wc -l             count lines ("word count", -l = lines only)
grep TEXT FILE    print only the lines that contain TEXT
sort -n           sort lines as numbers (-n); without -n it sorts as text
uniq -c           collapse repeated neighbouring lines, -c = with a count
awk '...'         a small language for columns: $1 is the first column,
                  $NF the last, NR the line number; {s+=$1} END {print s} sums
jq '...'          the same idea for JSON (a text format of {"key": value} records)

Two pieces of glue:

That is enough to start. Each mission says what the commands it needs are for; the hints get more concrete, and solution shows a full worked session. You will learn all of these properly in the Linux chapters.

Later (Ch 7): grep, sort, uniq, awk and jq get a whole chapter; here you only need the one-liners shown.

Doing it with the tools you have

This box's nginx writes the request duration ($request_time, in seconds) as the last column of every access-log line, and the status code as column 9. So the access log already contains three of the four signals:

# traffic: count the lines (one line = one request) and look at the first and last time
head -1 access.log; tail -1 access.log; wc -l < access.log

# errors: share of 5xx. For each line whose column 9 is 500 or more, add one to e;
# at the END print e as a percentage of NR (the number of lines read)
awk '$9 >= 500 {e++} END {printf "%.1f%%\n", 100*e/NR}' access.log

# latency: print every duration in ms, sort them as numbers, then pick the one at rank 99%
awk '{print $NF*1000}' access.log | sort -n |
  awk '{a[NR]=$1} END {i=int(NR*0.99); if (i<NR*0.99) i++; print "p99", a[i]}'

(; separates commands on one line; < access.log feeds the file to wc as input so it prints only the number; printf prints with a format, %.1f = one decimal, %% = a literal percent sign.)

The last pipeline sorts the values and picks the value at position 99% of the way down the list, rounded up: that is the nearest-rank percentile. Monitoring tools estimate percentiles slightly differently, so their numbers differ a little; the idea is the same.

What you can now do:

Why it helps

When you are paged for checkout and open the dashboard, these four are what should be on the first row, and you'll build that row. Situations: a manager says "average latency is 322 ms, fine"; you show p99 is 3 seconds because the pool times out. A service is at 15% CPU and failing; saturation, meaning requests waiting for a database connection, explains it and a CPU alert never would. An interviewer asks "what would you monitor for a new API?"; the golden signals plus RED and USE is the expected structure. And you'll resist the urge to put forty graphs on the top-level dashboard, because everything beyond these four is a cause you look at during diagnosis.

FAQ

Why percentiles instead of averages for latency?

An average describes nobody. 99 requests at 50 ms and one at 60 seconds average about 650 ms, a time no request experienced, and averages hide bimodal distributions like fast cache hits plus slow misses. Percentiles describe real users: p50 the median, p99 the one in a hundred who waited longest. And with fan-out (one page calling many services), p99 matters more than it looks: a page making 100 calls hits at least one p99 63% of the time.

Why keep latency of errors separate from latency of successes?

They are different populations that pull the numbers in opposite directions. A fast failure, a 503 in 2 ms from a safety switch that refuses requests while a dependency is down, makes mixed latency look better the more you fail. A slow failure, a 5-second timeout, makes it look worse and hides that successes are fine. Track success latency as the user-experience metric, and look at error latency to learn how things fail: timeouts versus rejections.

What exactly is saturation for a web service?

The resource that runs out first, which is rarely the CPU. Typical candidates: the database connection pool (connections in use versus the maximum, and requests waiting for one), the worker threads that each handle one request (busy versus maximum), the program's memory, file descriptors (the handles a program needs for every open file or connection) against their limit, and a CPU cap. A service can be at 15% CPU and completely saturated, with every request waiting for a connection.

What are RED and USE, and how do they relate to the golden signals?

RED is Rate, Errors, Duration, per service: the golden signals without saturation, used to see that a service hurts. USE is Utilisation, Saturation, Errors, per resource (CPU, disk, connection pool, network card), used to find which resource is the reason. You watch RED or the golden signals all the time, and walk USE during diagnosis.

Can I average the p99 of each server to get the service p99?

No. Percentiles do not average, because each one ignores how many requests it describes. A server with 10 slow requests would count as much as one with 1000 fast ones. You need the underlying distribution: raw values, or histogram bucket counts (how many requests fell under each of a few fixed durations) added up across servers, which is what monitoring systems compute their percentiles from.

In an interview Junior

What are the four golden signals?

They are the four things to measure first on a user-facing service, from the Google SRE book: latency, traffic, errors and saturation.

The first three describe what already happened. Saturation is the leading indicator: it moves before latency and errors fall off a cliff. From an nginx access log you get the first three with wc -l, awk '$9 >= 500' and sort -n; saturation needs the service's own numbers.

Also asked: Why is the average latency a bad number to put on a dashboard? · What does p99 mean, and why look at it as well as p50? · What is the difference between RED and USE?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.