OnCallReady

Lesson 30.4 · AWS II: VPC, EC2, ELB & EKS · 15 min read

Security groups vs network ACLs

In plain words

Two kinds of guards protect your machines. The security group is a doorman at each machine's own door with a guest list: once a guest is let in, the doorman also lets the guest's replies out without checking again. The network ACL is a checkpoint at the street corner with a numbered rule sheet: it checks every packet that passes, both ways, and it does not remember who came in a minute ago.

So the corner checkpoint needs a rule for the answers too, or the answers never leave the street.

Security groups vs network ACLs

Two firewalls sit between every pair of things in a VPC, and they behave differently enough that mixing them up causes a whole class of outages: "the rule allows port 8080, so it cannot be the firewall". This lesson is the difference, how to read each one, and the flow logs that show which one dropped a packet.

Need to know: a security group (SG) is a stateful allow-list attached to network interfaces: allow the request in and the reply goes out by itself; there are no deny rules; a rule's source can be another security group. A network ACL (NACL) is a stateless rule list on a subnet: numbered rules, allow and deny, evaluated lowest number first, and the reply is a separate packet that needs its own rule - to the client's ephemeral port (1024-65535).

Security groups

attached tonetwork interfaces (instances, load balancer nodes, endpoints, pods with security groups for pods)
rulesallow only; inbound and outbound lists
statestateful: the reply to an allowed connection is always allowed
source / destinationa CIDR, a prefix list, or another security group
defaulta new group allows all outbound, nothing inbound
limits60 inbound + 60 outbound rules per group, 5 groups per interface (defaults)

The lesson's VPC has the classic pair - the load balancer's group and the app's group, the app allowing 8080 from the load balancer's group:

$ VPC=$(aws ec2 describe-vpcs --filters Name=tag:Name,Values=try-vpc --query 'Vpcs[0].VpcId' --output text)
$ aws ec2 describe-security-groups --filters Name=vpc-id,Values=$VPC --query 'SecurityGroups[].[GroupName,GroupId,Description]' --output table
-------------------------------------------------------------------
|                     DescribeSecurityGroups                      |
+---------+------------------------+------------------------------+
|  default|  sg-093bab5539b6bb729  |  default VPC security group  |
|  try-alb|  sg-05441f878e288a142  |  try: the load balancer      |
|  try-app|  sg-0469a55e27b3c238a  |  try: the app servers        |
+---------+------------------------+------------------------------+
$ APP=$(aws ec2 describe-security-groups --filters Name=vpc-id,Values=$VPC Name=group-name,Values=try-app --query 'SecurityGroups[0].GroupId' --output text)
$ aws ec2 describe-security-group-rules --filters Name=group-id,Values=$APP --query 'SecurityGroupRules[].[IsEgress,IpProtocol,FromPort,ToPort,CidrIpv4,ReferencedGroupInfo.GroupId]' --output table
-----------------------------------------------------------------------
|                     DescribeSecurityGroupRules                      |
+-------+------+-------+-------+------------+-------------------------+
|  True |  -1  |  -1   |  -1   |  0.0.0.0/0 |  None                   |
|  False|  tcp |  8080 |  8080 |  None      |  sg-05441f878e288a142   |
+-------+------+-------+-------+------------+-------------------------+

Read the rows: inbound (IsEgress False) TCP 8080 from the group try-alb - whatever network interface carries that group, today or after the load balancer adds nodes; outbound (True) protocol -1 (all) to 0.0.0.0/0 - the default every new group gets.

Group references are the reason SGs scale: you never write an ALB's node IPs into a rule. The same works within a group: a rule "from this group itself" lets every member talk to every other member (the EKS cluster security group does exactly that).

Adding and removing rules is two calls, and AWS refuses duplicates:

$ aws ec2 authorize-security-group-ingress --group-id $APP --protocol tcp --port 9100 --cidr 10.99.0.0/16 --query 'SecurityGroupRules[0].SecurityGroupRuleId' --output text
sgr-05851770baea5ab83
$ aws ec2 authorize-security-group-ingress --group-id $APP --protocol tcp --port 9100 --cidr 10.99.0.0/16
aws: [ERROR]: An error occurred (InvalidPermission.Duplicate) when calling the AuthorizeSecurityGroupIngress operation: the specified rule "peer: 10.99.0.0/16, TCP, from port: 9100, to port: 9100, ALLOW" already exists
$ aws ec2 revoke-security-group-ingress --group-id $APP --protocol tcp --port 9100 --cidr 10.99.0.0/16
{
    "Return": true,
    "RevokedSecurityGroupRules": [
        {
            "SecurityGroupRuleId": "sgr-05851770baea5ab83",
            "GroupId": "sg-0469a55e27b3c238a",
            "IsEgress": false,
            "IpProtocol": "tcp",
            "FromPort": 9100,
            "ToPort": 9100,
            "CidrIpv4": "10.99.0.0/16"
        }
    ]
}

A blocked connection is dropped, never refused: the client waits for its timeout (curl exit 28), it does not get "connection refused". "Refused" means the packet arrived and nothing listens on the port - a different problem, on the instance.

In an interview: "Security group or NACL - what is the difference?" - SGs are stateful, allow-only and attached to interfaces, with other groups as sources; NACLs are stateless, numbered allow/deny rules on subnets, so the return traffic to ephemeral ports needs its own rule.

Network ACLs

Every subnet has exactly one NACL. The VPC's default NACL allows everything both ways, which is why most teams never touch NACLs. A custom NACL starts with only the final rule *: deny all.

$ aws ec2 describe-network-acls --filters Name=vpc-id,Values=$VPC Name=tag:Name,Values=try-private --query 'NetworkAcls[0].Entries[].[Egress,RuleNumber,Protocol,PortRange.From,PortRange.To,CidrBlock,RuleAction]' --output table
--------------------------------------------------------------------
|                        DescribeNetworkAcls                       |
+-------+--------+-----+-------+--------+----------------+---------+
|  True |  100   |  6  |  443  |  443   |  0.0.0.0/0     |  allow  |
|  True |  110   |  6  |  1024 |  65535 |  10.99.0.0/16  |  allow  |
|  True |  32767 |  -1 |  None |  None  |  0.0.0.0/0     |  deny   |
|  False|  100   |  6  |  8080 |  8080  |  10.99.0.0/16  |  allow  |
|  False|  110   |  6  |  1024 |  65535 |  0.0.0.0/0     |  allow  |
|  False|  32767 |  -1 |  None |  None  |  0.0.0.0/0     |  deny   |
+-------+--------+-----+-------+--------+----------------+---------+

Protocol is a number here (6 = TCP, 17 = UDP, -1 = all). Walk a request through it - a health check from the ALB (in a public subnet, 10.99.0.x) to the app on 8080:

  1. In to the private subnet, TCP to port 8080 from 10.99.0.x: rule 100 allows.
  2. The app answers from port 8080 to the ALB's ephemeral port, say 41234. Out, TCP to 41234: rule 100 (443) does not match, rule 110 (1024-65535 to the VPC) allows. Without rule 110, the * deny drops the reply - and the health check times out.
  3. The app calling an AWS API over HTTPS: out to 443 (rule 100), and the reply comes in to the app's ephemeral port: inbound rule 110.

Rules are evaluated in number order and the first match wins, so a deny at 90 beats an allow at 100 - the one thing NACLs can do that SGs cannot: block a specific address range in front of a whole subnet. Leave gaps between numbers (100, 110, 120) for the rule you will need to squeeze in.

Which one when

NeedUse
"only the load balancer may reach the app"SG with the ALB's SG as source
"the database only from the app tier"SG referencing the app's SG
block an abusive /24 for a whole subnet nowNACL deny rule with a low number
a compliance rule "subnet X never talks to subnet Y"NACL, as a second layer
anything elseSG; leave NACLs at the default allow-all

Azure's NSG is a bit of both: stateful like an SG, but with priorities and deny rules like a NACL, and attachable to a subnet or a NIC.

Seeing the drops: VPC flow logs

A flow log records the IP traffic of a VPC, a subnet or one interface, aggregated per 60 s or 600 s window, delivered to CloudWatch Logs or S3 (gzip). The default format, one line per flow:

version account-id interface-id srcaddr dstaddr srcport dstport protocol packets bytes start end action log-status
2 111122223333 eni-0a1b2c3d4e5f60789 10.99.10.21 10.99.0.37 8080 41234 6 2 120 1790106903 1790106960 REJECT OK

That line is the NACL story above: the app (.10.21, port 8080) replying to the ALB node's ephemeral port 41234, REJECTed on the way out. Pairs to remember: an ACCEPT in and a REJECT out for the same flow = a NACL missing the return rule (SGs never do that: they are stateful); nothing at all = the packet never reached the interface (routing, or the source's own rules). AWS's Reachability Analyzer answers "can A reach B on port P" from the configuration alone, without sending a packet.

What you can do now

Why it helps

"The firewall allows the port, so it cannot be the firewall" is one of the most expensive sentences in an outage. Knowing that security groups are stateful and network ACLs are not, that NACL replies go to ephemeral ports, and how to read a flow log line lets you find the dropped packet in minutes instead of hours, and design rules that are tight without breaking replies.

Commands in this lesson

aws

FAQ

Can a security group deny traffic?

No. Security groups only allow; anything not allowed is dropped. To block a specific address range you use a network ACL deny rule with a low rule number, which applies to the whole subnet. In practice most teams express everything with security groups and leave NACLs at the default allow-all.

What does it mean that a rule's source is another security group?

It allows traffic from any network interface that carries that group, whatever its IP. "8080 from the load balancer's group" keeps working when the load balancer adds or replaces nodes, and nobody has to maintain address lists. A group can also reference itself, letting all members talk to each other.

What are ephemeral ports?

The temporary source port a client picks for each connection, somewhere in 1024-65535 (Linux uses 32768-60999). The reply is sent back to that port. Because network ACLs are stateless, the reply needs its own rule: outbound to 1024-65535 on the server's subnet, inbound on the client's. Security groups handle this automatically.

Why did my curl hang instead of failing quickly?

Security groups and NACLs drop packets silently; they never send a refusal. The client waits until its timeout (curl exit code 28). "Connection refused" is different: the packet arrived and nothing listened on that port, which points at the application or the port number, not the firewall.

How do flow logs help?

They record each flow per network interface, subnet or VPC with source, destination, ports, protocol, packets, bytes and ACCEPT or REJECT. An ACCEPT in followed by a REJECT out for the same flow is the signature of a NACL missing its return rule; no record at all means the packet never reached the interface.

In an interview Mid

Security group or network ACL - what is the difference?

A security group is stateful and attached to network interfaces: allow rules only, the reply to an allowed connection always passes, and a rule's source can be another security group. A network ACL is stateless and attached to a subnet: numbered allow and deny rules evaluated lowest first, a final * deny, and the reply is a separate packet that needs its own rule to the client's ephemeral ports (1024-65535). Use security groups for almost everything and NACLs as a coarse extra layer.

Also asked: Why would an ALB health check time out when the security group allows the port? · How do you read a VPC flow log line? · When would you add a deny rule to a network ACL?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.