OnCallReady

Lesson 8.9 · Addressing & DNS · 16 min read

Cloud subnets and sizing a cluster of machines

In plain words

Imagine renting a car park. The owner always keeps five spaces for himself: the entrance, his own office, two for the security guard, and the exit. So a 64-space car park only gives you 59.

Now imagine each car you park brings trailers that need their own spaces. With one kind of rental (the flat design), every car reserves room for 30 trailers in your car park up front, parked or not. With another (the overlay design), the trailers go in a separate, private field, and your car park only holds the cars.

That is cluster sizing: the car park is the cloud subnet, the cars are the machines (nodes), the trailers are the apps on them, and the design decides whether the apps eat your subnet.

Why this matters

A common ticket: "we need a subnet for a group of 50 machines that each run 30 small apps - how big?". If the answer is too small, nothing breaks on day one. It breaks months later during an upgrade, at the worst moment, and a subnet in use cannot simply be made bigger. This lesson gives you the formula and the questions to ask first.

What you need to know already: 8.3 (sizes, usable hosts) and 8.6 (plans, alignment).

A few words first

The cloud keeps five addresses per subnet, not two

Cloud providers reserve more addresses in every subnet than a normal network does. The big providers take five (this layout is Microsoft's):

x.x.x.0     network address
x.x.x.1     the default gateway (the provider's virtual router)
x.x.x.2     \  used by the provider's DNS service
x.x.x.3     /
x.x.x.255   broadcast (the last address of the subnet, whatever the size)

So on such a cloud:

usable = 2^(32 - prefix) - 5
/29 -> 8 - 5  = 3       (usually the smallest subnet allowed)
/28 -> 16 - 5 = 11
/27 -> 32 - 5 = 27
/26 -> 64 - 5 = 59
/24 -> 256 - 5 = 251
/22 -> 1024 - 5 = 1019

Your first VM in 10.10.1.0/24 gets 10.10.1.4, not .1. That surprises everyone once.

Two more facts that bite:

Two ways to give addresses to the apps on a node

Picture each node running many small apps side by side, each app instance with its own IP address. Where do those addresses come from? There are two designs.

design     app addresses come from                 the subnet must fit
---------  --------------------------------------  -------------------------
flat       the SAME subnet as the nodes, reserved  nodes + nodes x apps
           up front: each node grabs a block for    per node
           its maximum number of apps
overlay    a separate private range that only the  nodes only
           cluster uses (for example a /24 per
           node out of 100.64.0.0/16)

The sizing formula (flat design)

addresses = (nodes + surge) + (nodes + surge) x apps_per_node

Worked, the classic question - 50 nodes, 30 apps each, flat, surge 1:

(50 + 1) + (50 + 1) x 30 = 51 + 1530 = 1581 addresses
/22 = 1024 - 5 = 1019   too small
/21 = 2048 - 5 = 2043   fits, with ~460 to spare
answer: /21

Same cluster, but it will grow to 60 nodes and upgrades use surge 3:

(60 + 3) + (60 + 3) x 30 = 63 + 1890 = 1953
/21 = 2043 usable       fits, barely - I would ask for a /20

And with 110 apps per node allowed instead of 30:

(50 + 1) + 51 x 110 = 5661   -> /19 (8187 usable)

A 4x jump from one setting. That is why "how many apps per node, at most?" is the first question to ask when someone hands you a subnet request.

The same cluster, overlay design

The subnet only needs the nodes (plus surge and load balancers):

nodes: 50 + 1 surge = 51  ->  /26 = 59 usable fits, /25 gives room to grow
apps:  each node gets a /24 from the private app range (256 addresses)
       100.64.0.0/16 holds 256 /24s -> at most 256 nodes with that range

Overlay moves the limit from "how much of the real network did we get" to "how big is the private app range" - and that range is private, so you can make it big.

How to answer the sizing question

Say the formula, state your assumptions out loud (flat or overlay, apps per node, surge, growth), do the arithmetic, round up to the next prefix, and mention the two traps: the provider's five reserved addresses, and that a subnet cannot be resized once it is in use. That reasoning is what the interviewer is checking.

Later (Ch 23): this is exactly the sizing of an Azure Kubernetes cluster; "flat" and "overlay" are its two network modes.

What you can now do

Why it helps

Subnet sizing is a very common platform ticket: "we need a subnet for a new cluster, how big?" Get it wrong and the failure comes later and at the worst moment: an upgrade fails because the extra machine it adds has no address left, or the cluster cannot grow during a traffic peak. And you cannot simply grow a subnet that is in use.

Knowing the formula lets you answer that ticket in a minute and ask the right questions first: flat or overlay, how many apps per machine at most, how many machines at peak, how many spare machines during upgrades. It also lets you spot a 4x jump when someone raises the apps-per-machine limit. The 50-machines-30-apps sizing question is a classic in platform interviews, and the reasoning matters more than the number.

FAQ

Why does my first machine in a new cloud subnet get .4 and not .1?

The big cloud providers reserve the first four addresses and the last one in every subnet: .0 is the network, .1 is the default gateway (the provider's virtual router), .2 and .3 are used for the provider's DNS service, and the last is broadcast. So usable = size minus 5, and the first address you can use is .4. This is on top of the normal rule, not instead of it.

Flat or overlay: which should a new cluster use?

Overlay, unless something outside the cluster must see each app's real address. Overlay keeps app addresses in a private range, so the subnet only needs the machines and you can make the private range big. Flat is chosen when outside firewalls or monitoring need a distinct real address per app, and you pay for it in subnet space.

Can I change the apps-per-machine limit later?

Usually not on machines that already exist: it is fixed when they are set up. You add a new group of machines with the new limit, move the apps over, and remove the old group. On a flat design that new group needs its own addresses while both groups exist, so plan the space for the move, not just for the end state.

Why does surge matter in the formula?

During an upgrade a new machine is added before an old one is removed, so for a while there are extra machines. With a surge of 1 there is one extra; faster upgrades use more. On a flat design each extra machine needs its own address plus its full block of app addresses. If the subnet has no room, the upgrade fails, which is why upgrades are usually the first thing to break on a tight subnet.

What is the catch with overlay?

Apps talking to the outside world leave with their machine's address (the machine does NAT). Outside firewalls, database allow-lists and logs see machine addresses, not app addresses, so you cannot write a firewall rule for one app by IP. The private app range must also not overlap anything the apps need to reach. And with a /16 range and a /24 per machine, the cluster tops out at 256 machines.

In an interview Junior

Someone asks for a subnet for a cluster of 50 machines, each running up to 30 apps that get their own address. What size do you give them?

First, state the assumptions: flat design (app addresses come from the same subnet), surge of 1 node during upgrades, no growth.

Then mention the traps: a subnet in use cannot be resized, so size for growth; load balancer front-ends take addresses too; and "how many apps per node, at most?" changes the answer the most. With an overlay design the subnet only needs the nodes, so a /26 or /25 would do.

Also asked: How many usable addresses does a /24 have on a cloud subnet, and why not 254? · Why can you not just make a subnet bigger later? · What is the difference between a flat and an overlay design for app addresses?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.