Reviewing an AWS workload: the network and compute checklist
You can now build every layer of a typical AWS workload. The skill that pays on-call is the reverse: looking at an account someone else built and seeing, in ten minutes, what will page you. This lesson is the checklist, in the order you walk it, with the commands that answer each line. It ends the chapter: once its labs are done, entering this lesson removes the instances, clusters and load balancers the chapter built (simulator), so the saved box stays small.
Need to know: review in the order traffic flows - network (CIDRs, AZs, routes), entry points (load balancers, DNS, what is public), compute (images, instance settings, groups and clusters), identity (roles of instances and pods, cluster access), data paths (how workloads reach AWS APIs and the internet). Then the bill: what costs money while doing nothing.
1. Network
| Question | Command |
|---|---|
| Which VPCs, which CIDRs - do any overlap? | aws ec2 describe-vpcs --query 'Vpcs[].[VpcId,CidrBlock,IsDefault]' |
| Two AZs for every tier? | describe-subnets grouped by Name/AZ |
| Which subnets are public (0.0.0.0/0 -> igw) and why? | describe-route-tables |
| Any blackhole routes? | describe-route-tables, routes in state blackhole |
| One NAT gateway per AZ? | describe-nat-gateways |
| Custom NACLs (and their return rules)? | describe-network-acls --filters Name=default,Values=false |
| Flow logs on? | describe-flow-logs |
$ aws ec2 describe-vpcs --query 'Vpcs[].[VpcId,CidrBlock,IsDefault,Tags[?Key==`Name`]|[0].Value]' --output text
vpc-03f40e92832b5247c 172.31.0.0/16 True None
$ aws ec2 describe-route-tables --query 'RouteTables[].Routes[?State==`blackhole`].[DestinationCidrBlock,NatGatewayId]' --output text
2. Entry points
- Every internet-facing load balancer: is it supposed to be public? HTTPS with a current certificate, HTTP only redirecting? (
describe-load-balancers,describe-listeners) - Target health right now, and the health check path - does it test the app or just "the process is up"?
- Security groups open to
0.0.0.0/0, especially on 22, 3389, databases:
$ aws ec2 describe-security-groups --filters Name=ip-permission.cidr,Values=0.0.0.0/0 --query 'SecurityGroups[].[GroupId,GroupName,VpcId]' --output text
- Instances with public IPs that are not load balancers or bastions:
$ aws ec2 describe-instances --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[?PublicIpAddress].[InstanceId,PublicIpAddress,Tags[?Key==`Name`]|[0].Value]' --output text
3. Compute
| Look for | Why |
|---|---|
HttpTokens: optional (IMDSv1 allowed) | the SSRF credential theft path |
| key pairs and port 22 | Session Manager makes them unnecessary |
| instances outside any Auto Scaling group | pets: who rebuilds them? |
| ASGs with EC2 health checks behind a load balancer | dead apps stay in service |
| AMIs older than the patch policy | describe-images CreationDate of what runs |
t instances under steady load | CPU credits run out: sudden slowness |
| EKS versions near the end of standard support | the extended-support price, forced upgrades |
$ aws ec2 describe-instances --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[].[InstanceId,MetadataOptions.HttpTokens,KeyName,InstanceType]' --output text
4. Identity
- Instance profiles: which role, what does it allow? Admin-like policies on instances are the most common finding.
- EKS: who has cluster-admin (access entries, aws-auth)? Do pods get their own roles (Pod Identity or IRSA), or the node role? Node hop limit 1 if they do not need IMDS.
- Trust policies of IRSA roles: exact
:sub, neversystem:serviceaccount:*.
5. Data paths
- S3 and DynamoDB traffic through a gateway endpoint (free) instead of the NAT gateway (per GB).
- Private-only VPCs: an endpoint for every API the workload calls.
- Endpoint policies where data exfiltration matters.
The bill: what costs money while idle
| Item | Roughly, eu-central-1 |
|---|---|
| NAT gateway | $0.052/h each (~$38/month) + $0.052/GB processed |
| Application / Network Load Balancer | ~$0.027/h (~$20/month) + capacity units |
| Interface VPC endpoint | ~$0.012/h per AZ each + $0.01/GB |
| EKS control plane | $0.10/h (~$73/month); $0.60/h on extended support |
| Public IPv4 address | $0.005/h each (~$3.6/month), in use or not |
| Unattached EBS volume, old snapshot | storage price, forever |
The two classic "why did the bill jump" answers in this chapter's territory: NAT gateway data processing (container image pulls, logs shipped to S3 without a gateway endpoint) and load balancers nobody deleted (one per Kubernetes Ingress without IngressGroups).
Tags make all of this answerable: team, env, service on everything (Terraform's default_tags), so a review can ask "whose is this NAT gateway" and Cost Explorer can split the bill
- the next chapter is about that bill and the monitoring around it.
In an interview: "You inherit an AWS account. What do you check first?" - walk it in traffic order: VPCs, routes and what is public; internet-facing load balancers and security groups open to the world; instances (IMDSv2, no keys, in Auto Scaling groups), EKS access and pod roles; how workloads reach AWS APIs; and the idle costs: NAT gateways, load balancers, endpoints, clusters.
Not covered here
Transit Gateway (many VPCs and VPNs in a hub), VPC peering, PrivateLink services of your own, CloudFront and Global Accelerator in front of load balancers, Route 53 Resolver for hybrid DNS, and the EKS add-ons for storage (EBS/EFS CSI) and autoscaling (Karpenter). Each builds on the pieces of this chapter.
What you can do now
- review an account's network, entry points, compute, identity and data paths with the CLI;
- point at what will page you and at what costs money doing nothing.