Why this matters
A common ticket: "we need a subnet for a group of 50 machines that each run 30 small apps - how big?". If the answer is too small, nothing breaks on day one. It breaks months later during an upgrade, at the worst moment, and a subnet in use cannot simply be made bigger. This lesson gives you the formula and the questions to ask first.
What you need to know already: 8.3 (sizes, usable hosts) and 8.6 (plans, alignment).
A few words first
- The cloud means renting machines and networks from a provider (the big ones are run by Microsoft, Amazon and Google) instead of owning them. You get virtual machines (VMs) - like your UTM lab VM, but in their data centre.
- A cloud network is a private address range you rent there, for example
10.60.0.0/16, which you split into subnets exactly like the last lesson. - A cluster is a group of machines that work together and share the work. Each machine in it is a node.
The cloud keeps five addresses per subnet, not two
Cloud providers reserve more addresses in every subnet than a normal network does. The big providers take five (this layout is Microsoft's):
x.x.x.0 network address
x.x.x.1 the default gateway (the provider's virtual router)
x.x.x.2 \ used by the provider's DNS service
x.x.x.3 /
x.x.x.255 broadcast (the last address of the subnet, whatever the size)
So on such a cloud:
usable = 2^(32 - prefix) - 5
/29 -> 8 - 5 = 3 (usually the smallest subnet allowed)
/28 -> 16 - 5 = 11
/27 -> 32 - 5 = 27
/26 -> 64 - 5 = 59
/24 -> 256 - 5 = 251
/22 -> 1024 - 5 = 1019
Your first VM in 10.10.1.0/24 gets 10.10.1.4, not .1. That surprises everyone once.
Two more facts that bite:
- You cannot resize a subnet that has machines in it (in practice, most providers refuse). Plan the size up front.
- Other things quietly take addresses from the same subnet: every load balancer front-end (a load balancer is a service that receives traffic on one address and spreads it over several machines), and every extra machine that exists for a while during an upgrade.
Two ways to give addresses to the apps on a node
Picture each node running many small apps side by side, each app instance with its own IP address. Where do those addresses come from? There are two designs.
design app addresses come from the subnet must fit
--------- -------------------------------------- -------------------------
flat the SAME subnet as the nodes, reserved nodes + nodes x apps
up front: each node grabs a block for per node
its maximum number of apps
overlay a separate private range that only the nodes only
cluster uses (for example a /24 per
node out of 100.64.0.0/16)
- Flat eats subnet space: every app gets a real address on the network, reserved per node whether the app exists yet or not. The upside: other networks (and firewalls) see each app's real address.
- Overlay keeps app addresses off the network. The subnet only holds the nodes. The price: when an app talks to the outside world, its traffic leaves with the node's address (the node does NAT, 8.6), so outside firewalls cannot tell the apps apart.
The sizing formula (flat design)
addresses = (nodes + surge) + (nodes + surge) x apps_per_node
- surge - an upgrade usually adds a fresh node before it retires an old one, so for a while you have extra nodes. One is common; faster upgrades use more. Every surge node needs its own address and its own block of app addresses.
- Use the node count you expect to grow to, not today's.
- Then add the provider's five, and room for load balancer front-ends.
Worked, the classic question - 50 nodes, 30 apps each, flat, surge 1:
(50 + 1) + (50 + 1) x 30 = 51 + 1530 = 1581 addresses
/22 = 1024 - 5 = 1019 too small
/21 = 2048 - 5 = 2043 fits, with ~460 to spare
answer: /21
Same cluster, but it will grow to 60 nodes and upgrades use surge 3:
(60 + 3) + (60 + 3) x 30 = 63 + 1890 = 1953
/21 = 2043 usable fits, barely - I would ask for a /20
And with 110 apps per node allowed instead of 30:
(50 + 1) + 51 x 110 = 5661 -> /19 (8187 usable)
A 4x jump from one setting. That is why "how many apps per node, at most?" is the first question to ask when someone hands you a subnet request.
The same cluster, overlay design
The subnet only needs the nodes (plus surge and load balancers):
nodes: 50 + 1 surge = 51 -> /26 = 59 usable fits, /25 gives room to grow
apps: each node gets a /24 from the private app range (256 addresses)
100.64.0.0/16 holds 256 /24s -> at most 256 nodes with that range
Overlay moves the limit from "how much of the real network did we get" to "how big is the private app range" - and that range is private, so you can make it big.
How to answer the sizing question
Say the formula, state your assumptions out loud (flat or overlay, apps per node, surge, growth), do the arithmetic, round up to the next prefix, and mention the two traps: the provider's five reserved addresses, and that a subnet cannot be resized once it is in use. That reasoning is what the interviewer is checking.
Later (Ch 23): this is exactly the sizing of an Azure Kubernetes cluster; "flat" and "overlay" are its two network modes.
What you can now do
- Count usable addresses on a cloud subnet (minus 5).
- Size a subnet for a cluster, flat or overlay, including surge and growth.
- Ask the right questions before agreeing to a size.