OnCallReady

Lesson 30.35 · AWS II: VPC, EC2, ELB & EKS · 10 min read

Reviewing an AWS workload: the network and compute checklist

In plain words

Reviewing an AWS setup is like inspecting a building before you take over as caretaker: you walk the same path a visitor walks - from the street, through the front door, up the stairs, into the offices - and at each point ask who can get in, what would break if this failed, and what keeps costing money with the lights off.

Reviewing an AWS workload: the network and compute checklist

You can now build every layer of a typical AWS workload. The skill that pays on-call is the reverse: looking at an account someone else built and seeing, in ten minutes, what will page you. This lesson is the checklist, in the order you walk it, with the commands that answer each line. It ends the chapter: once its labs are done, entering this lesson removes the instances, clusters and load balancers the chapter built (simulator), so the saved box stays small.

Need to know: review in the order traffic flows - network (CIDRs, AZs, routes), entry points (load balancers, DNS, what is public), compute (images, instance settings, groups and clusters), identity (roles of instances and pods, cluster access), data paths (how workloads reach AWS APIs and the internet). Then the bill: what costs money while doing nothing.

1. Network

QuestionCommand
Which VPCs, which CIDRs - do any overlap?aws ec2 describe-vpcs --query 'Vpcs[].[VpcId,CidrBlock,IsDefault]'
Two AZs for every tier?describe-subnets grouped by Name/AZ
Which subnets are public (0.0.0.0/0 -> igw) and why?describe-route-tables
Any blackhole routes?describe-route-tables, routes in state blackhole
One NAT gateway per AZ?describe-nat-gateways
Custom NACLs (and their return rules)?describe-network-acls --filters Name=default,Values=false
Flow logs on?describe-flow-logs
$ aws ec2 describe-vpcs --query 'Vpcs[].[VpcId,CidrBlock,IsDefault,Tags[?Key==`Name`]|[0].Value]' --output text
vpc-03f40e92832b5247c	172.31.0.0/16	True	None
$ aws ec2 describe-route-tables --query 'RouteTables[].Routes[?State==`blackhole`].[DestinationCidrBlock,NatGatewayId]' --output text

2. Entry points

$ aws ec2 describe-security-groups --filters Name=ip-permission.cidr,Values=0.0.0.0/0 --query 'SecurityGroups[].[GroupId,GroupName,VpcId]' --output text
$ aws ec2 describe-instances --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[?PublicIpAddress].[InstanceId,PublicIpAddress,Tags[?Key==`Name`]|[0].Value]' --output text

3. Compute

Look forWhy
HttpTokens: optional (IMDSv1 allowed)the SSRF credential theft path
key pairs and port 22Session Manager makes them unnecessary
instances outside any Auto Scaling grouppets: who rebuilds them?
ASGs with EC2 health checks behind a load balancerdead apps stay in service
AMIs older than the patch policydescribe-images CreationDate of what runs
t instances under steady loadCPU credits run out: sudden slowness
EKS versions near the end of standard supportthe extended-support price, forced upgrades
$ aws ec2 describe-instances --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[].[InstanceId,MetadataOptions.HttpTokens,KeyName,InstanceType]' --output text

4. Identity

5. Data paths

The bill: what costs money while idle

ItemRoughly, eu-central-1
NAT gateway$0.052/h each (~$38/month) + $0.052/GB processed
Application / Network Load Balancer~$0.027/h (~$20/month) + capacity units
Interface VPC endpoint~$0.012/h per AZ each + $0.01/GB
EKS control plane$0.10/h (~$73/month); $0.60/h on extended support
Public IPv4 address$0.005/h each (~$3.6/month), in use or not
Unattached EBS volume, old snapshotstorage price, forever

The two classic "why did the bill jump" answers in this chapter's territory: NAT gateway data processing (container image pulls, logs shipped to S3 without a gateway endpoint) and load balancers nobody deleted (one per Kubernetes Ingress without IngressGroups).

Tags make all of this answerable: team, env, service on everything (Terraform's default_tags), so a review can ask "whose is this NAT gateway" and Cost Explorer can split the bill

In an interview: "You inherit an AWS account. What do you check first?" - walk it in traffic order: VPCs, routes and what is public; internet-facing load balancers and security groups open to the world; instances (IMDSv2, no keys, in Auto Scaling groups), EKS access and pod roles; how workloads reach AWS APIs; and the idle costs: NAT gateways, load balancers, endpoints, clusters.

Not covered here

Transit Gateway (many VPCs and VPNs in a hub), VPC peering, PrivateLink services of your own, CloudFront and Global Accelerator in front of load balancers, Route 53 Resolver for hybrid DNS, and the EKS add-ons for storage (EBS/EFS CSI) and autoscaling (Karpenter). Each builds on the pieces of this chapter.

What you can do now

Why it helps

On a new team you inherit accounts you did not build. A structured walk - network, entry points, compute, identity, data paths, idle costs - finds the things that will page you (a single NAT gateway, IMDSv1, open security groups, pods on the node role) before they do, and gives you a list to fix in order. It is also exactly what interviewers ask a platform engineer to do.

Commands in this lesson

aws

FAQ

What should I look at first in an unknown account?

The network and what is public: VPCs and their CIDRs, which subnets route to an internet gateway, internet-facing load balancers, instances with public IPs and security groups open to 0.0.0.0/0. Exposure problems are the most urgent; reliability and cost come next.

Which AWS resources cost money while idle?

NAT gateways (per hour and per GB), load balancers, interface VPC endpoints, EKS control planes (six times more on extended support), public IPv4 addresses, unattached EBS volumes and old snapshots. They keep billing whether or not any traffic flows, and they are easy to forget after a test.

What compute settings are common findings?

IMDSv1 still allowed, key pairs and port 22 open, instances outside any Auto Scaling group, groups behind a load balancer using only EC2 health checks, old AMIs, burstable t instances under steady load, and EKS clusters close to the end of standard support.

Why do tags matter in a review?

Without team, environment and service tags you cannot answer whose a NAT gateway or a load balancer is, whether it is still needed, or split the bill. Terraform's default_tags make consistent tagging cheap. Cost Explorer and the next chapter's budgets depend on them.

What did this chapter not cover?

Transit Gateway and VPC peering for connecting many networks, PrivateLink services of your own, CloudFront and Global Accelerator in front of load balancers, Route 53 Resolver for hybrid DNS, and EKS storage and autoscaling add-ons such as the EBS CSI driver and Karpenter. They build on the same pieces.

In an interview Mid

You inherit an AWS account. What do you check first?

Walk it in the order traffic flows: the network (VPCs, CIDR overlaps, which subnets are public, blackhole routes, one NAT gateway per AZ), the entry points (internet-facing load balancers, security groups open to 0.0.0.0/0, instances with public IPs), the compute (IMDSv2 required, no keys, instances in Auto Scaling groups, EKS versions), the identity (instance and pod roles, who is cluster-admin), the data paths (S3 gateway endpoints) - and the idle costs: NAT gateways, load balancers, endpoints, clusters.

Also asked: How would you find all security groups open to the internet? · What are common causes of a sudden jump in an AWS bill? · How do you make an AWS account easier to review next time?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.