OnCallReady

Lesson 30.7 · AWS II: VPC, EC2, ELB & EKS · 11 min read

The way out: NAT gateways and VPC endpoints

In plain words

Private machines are like staff in an office without windows: nobody outside can call them, but they still need to make phone calls out. A NAT gateway is the switchboard by the front door: staff dial out through it, and callers outside only ever see the switchboard's number, never the staff's. VPC endpoints are private internal lines straight to AWS's own departments, so staff can reach S3 or SSM without going outside at all.

The way out: NAT gateways and VPC endpoints

Private subnets exist so that nothing on the internet can start a connection to what lives there. But what lives there still needs to go out: package mirrors, container registries, a payment provider's API, and above all AWS's own APIs - S3, SSM, STS, ECR, CloudWatch. Two mechanisms provide the way out, and they differ in cost, in scope and in how they fail.

Need to know: a NAT gateway sits in a public subnet with an Elastic IP; private route tables send 0.0.0.0/0 to it, and it translates outbound connections to its public address (nothing can connect in through it). VPC endpoints reach AWS services without the internet: a gateway endpoint (S3 and DynamoDB only) is a route to a prefix list, free; an interface endpoint (most other services) is a network interface in your subnets with a security group and private DNS, billed per hour and per GB.

NAT gateways

$ aws ec2 describe-nat-gateways --filter Name=tag:Name,Values=try-nat --query 'NatGateways[].[NatGatewayId,State,SubnetId,ConnectivityType,NatGatewayAddresses[0].PublicIp,NatGatewayAddresses[0].PrivateIp]' --output table
------------------------------------------------------------------------------------------------------------
|                                            DescribeNatGateways                                           |
+-----------------------+------------+----------------------------+---------+---------------+--------------+
|  nat-07bc87b57a6297226|  available |  subnet-07d93f499311787f9  |  public |  3.66.139.168 |  10.99.0.31  |
+-----------------------+------------+----------------------------+---------+---------------+--------------+
$ aws ec2 describe-route-tables --filters Name=tag:Name,Values=try-private --query 'RouteTables[0].Routes[].[DestinationCidrBlock,NatGatewayId,GatewayId,State]' --output text
10.99.0.0/16	None	local	active
0.0.0.0/0	nat-07bc87b57a6297226	None	active

The private route table's default route points at the NAT gateway, which lives in try-public-a and has a public and a private address. Three rules about NAT gateways:

  1. It must be in a public subnet. Its own traffic leaves through that subnet's route to the IGW. A NAT gateway in a private subnet is a very expensive way to reach nothing.
  2. It lives in one AZ. If eu-central-1a has a problem, every private subnet routed through a NAT gateway in 1a loses the internet. Production runs one NAT gateway per AZ, and each AZ's private route table uses its own.
  3. Deleting it leaves a blackhole. The route stays, pointing at a NAT gateway that no longer exists, in state blackhole. Nothing errors; everything just times out.

It also costs money just for existing: in eu-central-1 about $0.052 per hour (~$38 a month) per NAT gateway, plus $0.052 per GB processed - the GB part is what surprises people when a cluster pulls container images or ships logs through it.

Gateway endpoints: S3 and DynamoDB by route

A gateway endpoint adds a route for the service's prefix list (the AWS-managed list of the service's address ranges in the Region) to the route tables you choose. Traffic to S3 then takes that route instead of 0.0.0.0/0 - longest prefix wins. It costs nothing and removes S3 traffic from the NAT gateway's per-GB bill, which is why every VPC should have one.

$ aws ec2 describe-prefix-lists --filters Name=prefix-list-name,Values=com.amazonaws.eu-central-1.s3 --query 'PrefixLists[0].[PrefixListId,PrefixListName]' --output text
pl-6ea54007	com.amazonaws.eu-central-1.s3
$ VPC=$(aws ec2 describe-vpcs --filters Name=tag:Name,Values=try-vpc --query 'Vpcs[0].VpcId' --output text)
$ RT=$(aws ec2 describe-route-tables --filters Name=tag:Name,Values=try-private --query 'RouteTables[0].RouteTableId' --output text)
$ aws ec2 create-vpc-endpoint --vpc-id $VPC --vpc-endpoint-type Gateway --service-name com.amazonaws.eu-central-1.s3 --route-table-ids $RT --query 'VpcEndpoint.[VpcEndpointId,State]' --output text
vpce-0c3a57523c55b1bf7	available
$ aws ec2 describe-route-tables --route-table-ids $RT --query 'RouteTables[0].Routes[].[DestinationCidrBlock || DestinationPrefixListId,NatGatewayId || GatewayId,State]' --output text
10.99.0.0/16	local	active
0.0.0.0/0	nat-07bc87b57a6297226	active
pl-6ea54007	vpce-0c3a57523c55b1bf7	active

The new route's destination is pl-... and its target the vpce-.... Gateway endpoints work only from inside the VPC (not from a VPN or a peered VPC), and only for S3 and DynamoDB.

Interface endpoints: PrivateLink

For everything else - SSM, STS, ECR, CloudWatch Logs, KMS, Secrets Manager - an interface endpoint puts a network interface (with a private IP) into each subnet you pick. With private DNS on, the service's normal name (ssm.eu-central-1.amazonaws.com) resolves to those private IPs inside the VPC, so no client configuration changes. What it needs:

Price: about $0.012 per hour per AZ per endpoint plus $0.01 per GB. Three services in two AZs is ~$52 a month - cheaper than a NAT gateway only when the list stays short.

Both endpoint types take an endpoint policy (a resource policy on the endpoint itself): the classic use is "this VPC may only reach our S3 buckets", which stops an instance from exfiltrating data to an attacker's bucket over the endpoint.

A private-only VPC

Some workloads must have no internet path at all: no IGW, no NAT. They still work if every AWS API they call has an endpoint. The failure mode is quiet: a call to a service without an endpoint does not fail with an error, it hangs until the SDK's connect timeout:

$ aws sts get-caller-identity
aws: [ERROR]: Connect timeout on endpoint URL: "https://sts.eu-central-1.amazonaws.com/"

That message from inside a VPC always means: no route to that endpoint (no NAT, no interface endpoint for this service), or an endpoint whose security group does not allow you.

In an interview: "How does an instance in a private subnet reach S3?" - through a NAT gateway in a public subnet, or better through an S3 gateway endpoint (a free prefix-list route that keeps the traffic off the internet and off the NAT bill).

What you can do now

Why it helps

Private subnets that cannot reach out break quietly: package installs hang, Session Manager says "not connected", AWS SDK calls time out. Knowing where the way out is (a NAT gateway in a public subnet, endpoints per service), how a deleted NAT gateway leaves a blackhole route, and what each option costs lets you fix those outages fast and keep the NAT bill from surprising anyone.

Commands in this lesson

aws

FAQ

Why must a NAT gateway be in a public subnet?

Because it sends its own traffic out through its subnet's route to the internet gateway. Placed in a private subnet it has no way out, so everything routed through it fails. The private subnets' route tables point 0.0.0.0/0 at the NAT gateway; the NAT gateway's subnet points 0.0.0.0/0 at the internet gateway.

Do I need one NAT gateway per AZ?

For production, yes. A NAT gateway lives in one AZ; if that zone has a problem, every private subnet routed through it loses the internet, including those in healthy zones. One NAT gateway per AZ, each AZ's private route table using its own, keeps a zone failure contained. Dev environments often accept a single one to save money.

What is a blackhole route?

A route whose target no longer exists, typically after someone deleted a NAT gateway. The route stays in the route table with state blackhole and traffic to it is dropped without any error. Replacing the target with replace-route fixes it; create-route would fail because the destination already has a route.

Why use an S3 gateway endpoint if there is a NAT gateway?

Cost and privacy. NAT gateways charge per GB processed, and S3 traffic (backups, logs, container image layers from ECR) can be most of it. A gateway endpoint is a free prefix-list route that sends S3 traffic over the AWS network directly. Longest-prefix matching makes it win over the 0.0.0.0/0 route automatically.

What does "Connect timeout on endpoint URL" mean from an instance?

The CLI or SDK could not open a connection to the service's endpoint: there is no NAT route and no interface endpoint for that service, or the endpoint's security group does not allow the instance on 443. In a private-only VPC every AWS API the workload calls needs its own endpoint, STS included.

In an interview Mid

How does an instance in a private subnet reach S3?

Either through a NAT gateway in a public subnet (the private route table's 0.0.0.0/0 points at it, and the NAT gateway reaches out through the internet gateway), or better through a gateway endpoint for S3: a free route to the S3 prefix list on the private route table, which keeps the traffic on the AWS network and off the NAT gateway's per-GB bill. Other services need a NAT gateway or an interface endpoint each, with a security group allowing 443 and private DNS.

Also asked: What happens to private instances when their NAT gateway is deleted? · What is the difference between a gateway endpoint and an interface endpoint? · How would you keep instances in a VPC from uploading data to buckets outside your company?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.