OnCallReady

AWS II: VPC, EC2, ELB & EKS: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 30 of the course.

A user says your site on AWS is down. Where do you look, in what order? Mid

Follow the request: DNS (does the name resolve to the load balancer), the load balancer itself (describe-target-health: how many targets are healthy and the reason code), the status code (502 refused, 503 no target, 504 timeout), then the path to the targets - the targets' security group allowing the load balancer's group, the subnet's network ACL in both directions - and finally the instances or pods (Session Manager, cloud-init status, the app's logs). For private workloads that call out, check the route table for a NAT gateway or endpoint and blackhole routes.

Also asked: What is the difference between a security group and a network ACL? · How do instances in a private subnet reach the internet? · How does a pod on EKS get AWS permissions?

What makes a subnet public in AWS? Mid

Only its route table: a 0.0.0.0/0 route to an internet gateway. There is no public flag on the subnet itself. An instance there also needs a public IPv4 address, because the internet gateway only translates addresses that have one. Private subnets have no such route; they reach out through a NAT gateway in a public subnet, or not at all. Subnets that are not associated explicitly use the main route table, which should stay local-only.

Also asked: How do you plan the CIDR ranges for a new VPC? · Why should every tier have a subnet in at least two AZs? · What happens if two VPCs you want to connect have overlapping CIDRs?

Learn it: 30.1 VPC design: CIDRs, subnets per AZ, route tables

Security group or network ACL - what is the difference? Mid

A security group is stateful and attached to network interfaces: allow rules only, the reply to an allowed connection always passes, and a rule's source can be another security group. A network ACL is stateless and attached to a subnet: numbered allow and deny rules evaluated lowest first, a final * deny, and the reply is a separate packet that needs its own rule to the client's ephemeral ports (1024-65535). Use security groups for almost everything and NACLs as a coarse extra layer.

Also asked: Why would an ALB health check time out when the security group allows the port? · How do you read a VPC flow log line? · When would you add a deny rule to a network ACL?

Learn it: 30.4 Security groups vs network ACLs

How does an instance in a private subnet reach S3? Mid

Either through a NAT gateway in a public subnet (the private route table's 0.0.0.0/0 points at it, and the NAT gateway reaches out through the internet gateway), or better through a gateway endpoint for S3: a free route to the S3 prefix list on the private route table, which keeps the traffic on the AWS network and off the NAT gateway's per-GB bill. Other services need a NAT gateway or an interface endpoint each, with a security group allowing 443 and private DNS.

Also asked: What happens to private instances when their NAT gateway is deleted? · What is the difference between a gateway endpoint and an interface endpoint? · How would you keep instances in a VPC from uploading data to buckets outside your company?

Learn it: 30.7 The way out: NAT gateways and VPC endpoints

How do you store Terraform state for a team on AWS? Mid

In an S3 backend: a dedicated bucket with versioning, Block Public Access and encryption, one key per stack and environment, and locking with use_lockfile = true on Terraform 1.10+ (a .tflock object written with a conditional PUT; a DynamoDB table on older versions). CI is the only regular writer, people use plan locally, and terraform force-unlock is only for locks whose run is really gone.

Also asked: What happens when someone changes a Terraform-managed security group in the console? · Why should you not mix inline security group rules with separate rule resources? · How do you pass outputs from a network stack to an application stack?

Learn it: 30.12 AWS with Terraform: the aws provider and state in S3

The root disk of an EC2 instance is full. How do you grow it without downtime? Mid

Grow the EBS volume with aws ec2 modify-volume --size while the instance runs (only up, once per six hours; usable once the modification is optimizing). Then inside the instance grow the partition with growpart (disk, space, partition number) and the file system with resize2fs for ext4 or xfs_growfs for XFS, and confirm with df -h. lsblk shows which layer is still small. Then ask why the data is on the root volume.

Also asked: Where do you look when an instance did not do what its user data says? · What is the difference between gp2 and gp3 volumes? · Why use a launch template instead of run-instances parameters?

Learn it: 30.14 EC2: images, instance types, user data, EBS

How does an application on EC2 get AWS credentials without access keys? Mid

Through the instance profile: the role's temporary credentials are served by the instance metadata service at 169.254.169.254, rotated by AWS, and the SDK reads them at the end of its credential chain. With IMDSv2 the client first PUTs /latest/api/token and sends the token header on every request, which blocks the SSRF attack that stole metadata credentials in 2019; HttpTokens=required enforces it, and the hop limit decides whether containers can reach it.

Also asked: Why should IMDSv1 be disabled? · How does Session Manager work without an open port? · What would you check if an instance does not appear in Session Manager?

Learn it: 30.17 Inside the instance: IMDSv2, instance roles, Session Manager

What is the difference between 502, 503 and 504 from an ALB? Mid

502 Bad Gateway: the ALB reached a target that refused or broke the connection or sent an invalid response - a wrong port, a crashed app, a keep-alive timeout shorter than the ALB's. 503: there was no target to send to - no registered targets, or no rule forwarding. 504: the target did not answer in time - a security group or network ACL dropping packets, or an overloaded app. describe-target-health and its reason codes show which layer.

Also asked: How do ALB health checks decide a target is healthy? · When would you choose an NLB over an ALB? · What does fail-open mean for a target group?

Learn it: 30.19 Load balancers: ALB, NLB, health checks and their status codes

How do you roll out a new AMI to an Auto Scaling group without downtime? Mid

Create a new launch template version with the new image (create-launch-template-version --source-version keeps the rest), then start-instance-refresh with a MinHealthyPercentage and an InstanceWarmup: instances are replaced in batches, each new one has to pass the load balancer's health checks and warm up before the next batch, and capacity never drops below the minimum. With AutoRollback a refresh whose new instances never become healthy goes back to the previous version.

Also asked: What is the difference between min, max and desired capacity? · Why would an Auto Scaling group keep replacing healthy-looking instances? · How does target tracking scaling work?

Learn it: 30.24 Auto Scaling groups: desired capacity, health, instance refresh

What does AWS manage in EKS and what do you manage? Mid

AWS runs the control plane - API servers and etcd across three AZs - patches it and offers managed add-ons (vpc-cni, coredns, kube-proxy, eks-pod-identity-agent). You own the nodes (a managed node group is an Auto Scaling group of EKS-optimized AMIs; or Fargate/Auto Mode), the VPC and the IP budget the VPC CNI consumes, who may use the cluster (access entries), pod identities, add-on versions, upgrades before standard support ends, and the workloads. kubectl authenticates with IAM through aws eks get-token.

Also asked: Why does the VPC CNI matter for subnet sizing? · How do you upgrade an EKS cluster? · What is the difference between managed node groups and Fargate?

Learn it: 30.26 EKS: what AWS runs and what you run

How does a pod on EKS get AWS permissions without access keys? Mid

With EKS Pod Identity: the eks-pod-identity-agent add-on, a role trusted by pods.eks.amazonaws.com for sts:AssumeRole and sts:TagSession, and an association of namespace and service account to the role; new pods get credentials from the agent. Or with IRSA: the cluster's OIDC issuer registered in IAM, a role whose trust allows AssumeRoleWithWebIdentity for the exact :sub system:serviceaccount:<ns>:<sa> and :aud sts.amazonaws.com, and the role ARN annotation on the service account. Without either, pods fall back to the node's role.

Also asked: What is the difference between Unauthorized and Forbidden from kubectl on EKS? · What do access policies like AmazonEKSEditPolicy do? · What does "Not authorized to perform sts:AssumeRoleWithWebIdentity" tell you?

Learn it: 30.29 EKS identity: access entries, IRSA, Pod Identity

How do you expose a service on EKS with an AWS load balancer? Mid

Install the AWS Load Balancer Controller with an IAM role (Pod Identity or IRSA), tag the subnets kubernetes.io/role/elb (public) or internal-elb (private), then create an Ingress with ingressClassName: alb and annotations for the scheme and target-type: ip - the controller builds an ALB whose targets are the pod IPs. For TCP, a Service of type LoadBalancer with the NLB annotations. kubectl describe ingress shows the controller's events when it cannot build it.

Also asked: What is an IngressGroup and why use one? · Why would an Ingress stay without an address? · What is the difference between target-type ip and instance?

Learn it: 30.33 The AWS Load Balancer Controller: Ingress and Service to ALB and NLB

You inherit an AWS account. What do you check first? Mid

Walk it in the order traffic flows: the network (VPCs, CIDR overlaps, which subnets are public, blackhole routes, one NAT gateway per AZ), the entry points (internet-facing load balancers, security groups open to 0.0.0.0/0, instances with public IPs), the compute (IMDSv2 required, no keys, instances in Auto Scaling groups, EKS versions), the identity (instance and pod roles, who is cluster-admin), the data paths (S3 gateway endpoints) - and the idle costs: NAT gateways, load balancers, endpoints, clusters.

Also asked: How would you find all security groups open to the internet? · What are common causes of a sudden jump in an AWS bill? · How do you make an AWS account easier to review next time?

Learn it: 30.35 Reviewing an AWS workload: the network and compute checklist

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.