OnCallReady

Kubernetes: Networking & Storage: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 16 of the course.

Walk me through what happens when a pod calls http://web.shop. Junior

  1. DNS - the pod's /etc/resolv.conf points at 10.96.0.10, the kube-dns Service in front of the CoreDNS pods. With ndots:5 and the search list, web.shop becomes web.shop.svc.cluster.local, which CoreDNS answers from the API with the Service's ClusterIP.
  2. Connect - the pod opens TCP to ClusterIP:80. Nothing listens there: kube-proxy has written iptables rules on every node, so the node's kernel DNATs the first packet to one Ready pod IP and port from the Service's EndpointSlice, chosen per connection. conntrack remembers the choice.
  3. Route - the CNI plugin (Calico) routes the packet to the pod: on the same node through its veth, otherwise through the tunnel to the node that owns that pod block.
  4. Policy - if a NetworkPolicy selects the destination, the CNI plugin allows or drops it.
  5. The pod answers on its targetPort; replies are un-NATed on the way back.

No endpoints gives an instant "connection refused"; a dropped packet times out.

Also asked: What is a Kubernetes Service, and why do you need one? · How would you expose an HTTP application to users outside the cluster? · What is the difference between a PersistentVolume and a PersistentVolumeClaim?

What is a Kubernetes Service, and why do you need one? Junior

Pod IPs change every time a pod is replaced, so nothing can hard-code one. A Service gives a group of pods one stable address and DNS name, and a label selector that says which pods are behind it right now.

k expose deploy web --port=80 --target-port=http creates one. To troubleshoot, k describe svc web: check the Selector matches the pods' labels, and that Endpoints is not empty. Then connect to the Service port, not the container port.

Also asked: A developer says "the Service does not work". How do you troubleshoot it? · What is the difference between port and targetPort? · What is an EndpointSlice?

Learn it: 16.1 Services: a stable address in front of moving pods

How does kube-proxy implement a ClusterIP Service? Mid

Nothing listens on a ClusterIP; it exists only as packet-rewriting rules on every node, written by kube-proxy (a DaemonSet) from Services and EndpointSlices. In the default iptables mode:

  1. KUBE-SERVICES (in the nat table) matches ClusterIP + port and jumps to the Service's KUBE-SVC-... chain.
  2. That chain picks an endpoint by probability (1/3, then 1/2, then the rest for three endpoints).
  3. Each KUBE-SEP-... chain DNATs to one pod IP:port; traffic from outside the pod network is masqueraded (SNAT) so replies come back the same way.
  4. conntrack remembers the choice, so the decision is once per connection - a long-lived HTTP/2 or gRPC connection sticks to one pod.

kube-proxy is not in the data path: if it dies, the rules keep working, they just stop updating. A Service with no endpoints gets a REJECT rule instead - an instant "connection refused". Other modes: IPVS (kernel load balancer), nftables, or eBPF plugins that replace kube-proxy.

Also asked: Why does ping to a ClusterIP not work, while a TCP connection does? · Why can gRPC traffic all land on one pod behind a Service? · What does "connection refused" from a Service tell you, compared with a timeout?

Learn it: 16.3 How a ClusterIP really routes: kube-proxy, iptables, IPVS

Explain the Kubernetes Service types. Junior

Each type builds on the one before:

For many HTTP sites behind one address you add an Ingress or Gateway on top.

Also asked: What does externalTrafficPolicy: Local change? · How do you expose Services on a bare-metal cluster with no cloud provider? · When would you use a headless Service?

Learn it: 16.6 Service types: NodePort, LoadBalancer, ExternalName, headless

How does service discovery work in Kubernetes? Junior

Through DNS served by CoreDNS:

So web works from the same namespace, web.shop from anywhere. DNS is not a security boundary.

To debug: k run dnsdebug --rm -it --image=busybox:1.36 --restart=Never -- nslookup kubernetes.default. A timeout means CoreDNS is down or UDP/TCP 53 is blocked; NXDOMAIN means the name is wrong. Beware dig, which ignores the search list.

Also asked: All pods in a namespace fail to resolve names. How do you debug it? · What is the Corefile, and where does it live? · What is the difference between dnsPolicy ClusterFirst and Default?

Learn it: 16.13 CoreDNS and the names it serves

What does ndots:5 do in a pod's resolv.conf, and why can it be a problem? Mid

The ndots rule: a name with fewer dots than ndots is tried with every search domain appended first, and only then as written. Kubernetes sets ndots:5 so that web, web.shop and web.shop.svc resolve through the search list.

The cost lands on external names. For api.github.com (2 dots) from a pod in shop:

  1. api.github.com.shop.svc.cluster.local - NXDOMAIN
  2. api.github.com.svc.cluster.local - NXDOMAIN
  3. api.github.com.cluster.local - NXDOMAIN
  4. api.github.com - the answer

With A and AAAA asked together, that is 8 queries for one name, 6 wasted - load on CoreDNS and latency per connection. You can see it with CoreDNS's log plugin.

Fixes: a trailing dot (api.github.com., no search; test the client), a lower ndots for the pod through dnsConfig.options, NodeLocal DNSCache, and caching or connection reuse in the app.

Also asked: How would you check how many DNS queries a pod makes for an external name? · What is NodeLocal DNSCache, and what does it fix? · What does a trailing dot on a hostname do?

Learn it: 16.14 ndots:5, the search path, and why api.github.com costs 8 queries

What is the difference between an Ingress and an Ingress controller? Junior

The link is the IngressClass: ingressClassName: nginx says which controller serves it; an Ingress without a class is ignored unless a default class exists.

How it fails tells you where to look: an empty ADDRESS = no controller claimed it; 404 = no host/path rule matched (check pathType); 503 = matched, but the Service has no ready endpoints; 502 = the pod refused or broke the connection. The controller's access log has a line per request with the upstream it used.

Also asked: Users get a 503 from the ingress for one app while others work. How do you debug it? · What is the difference between pathType Exact and Prefix? · How do you test an Ingress for a hostname that does not resolve yet?

Learn it: 16.19 Ingress: the resource, the controller, the rules

How do you configure TLS for an application behind an Ingress? Junior

Terminate TLS at the Ingress controller:

  1. Put the certificate - leaf plus intermediates - and its private key in a Secret of type kubernetes.io/tls in the Ingress's namespace: k create secret tls shop-tls --cert=shop.lab.crt --key=shop.lab.key.
  2. Reference it in the Ingress: tls: - hosts: [shop.lab] secretName: shop-tls. The hosts must match the rules and the certificate's names.
  3. The controller picks the certificate per request by SNI, decrypts, routes on host and path, and talks HTTP to the pods (or HTTPS with an annotation). Plain HTTP gets a 308 redirect.

Check it: curl -sv --resolve shop.lab:443:IP https://shop.lab/ or openssl s_client -connect IP:443 -servername shop.lab. "Kubernetes Ingress Controller Fake Certificate" means the controller found no usable Secret for that name - wrong name, wrong type or missing; its log says which.

In production, cert-manager issues and renews the Secret from an Issuer; still alert on expiry.

Also asked: How would you prevent certificate-expiry outages on a Kubernetes platform? · What does the Ingress controller's Fake Certificate tell you? · What is TLS passthrough, and what do you lose with it?

Learn it: 16.22 TLS termination at the Ingress

What problems does Gateway API solve compared with Ingress? Mid

Ingress is one object for everything, so anything beyond host and path (weights, header matches, rewrites, timeouts) leaked into controller-specific annotations, and one team's Ingress could break a shared host. Gateway API splits the job across objects with different owners:

Two more wins: status conditions say exactly why something is not working (Accepted, ResolvedRefs, with reasons like NotAllowedByListeners), and a route to another namespace's Service needs that namespace's ReferenceGrant - the Service owner decides.

It comes as CRDs plus an implementation; with ingress-nginx retired, it is the default for new platforms.

Also asked: Name the main Gateway API resources and who typically owns each. · How do you do a weighted canary release with an HTTPRoute? · What is a ReferenceGrant for?

Learn it: 16.26 Gateway API: GatewayClass, Gateway, HTTPRoute

What is a NetworkPolicy, and what is the default behaviour without one? Junior

Without any policy, everything talks to everything: any pod can reach any pod in any namespace on any port. Namespaces are not a network boundary.

A NetworkPolicy is a firewall rule for pods: it selects pods by label and lists which traffic may reach them (ingress) and which they may send (egress). The rule people get wrong: a pod is isolated in a direction only when some policy selects it and lists that direction in policyTypes; then only what some rule allows passes. Policies only add allowances - there is no "deny".

The usual start in a namespace:

  1. Default deny - podSelector: {}, policyTypes: [Ingress, Egress].
  2. Allow DNS straight after - egress to the k8s-app=kube-dns pods in kube-system on UDP and TCP 53, or every name lookup breaks.
  3. Allow each needed connection explicitly.

Blocked traffic is dropped (a timeout), and the CNI plugin enforces it: Calico and Cilium do, Flannel silently ignores policies. Prove each with a connection that should fail.

Also asked: You apply a default-deny policy and the application breaks in unexpected ways. What do you check? · Why must a NetworkPolicy target CoreDNS pods rather than the DNS Service IP? · Which component actually enforces NetworkPolicies?

Learn it: 16.29 NetworkPolicy: default allow, isolation, and the rules

What is the difference between namespaceSelector and podSelector in one "from" item, versus in two items? Mid

YAML's leading - starts a new list item, and in a NetworkPolicy that one dash changes the meaning:

- from:
  - namespaceSelector: {matchLabels: {team: edge}}
    podSelector: {matchLabels: {app: gateway}}

One item: AND - only app=gateway pods in namespaces labelled team=edge.

- from:
  - namespaceSelector: {matchLabels: {team: edge}}
  - podSelector: {matchLabels: {app: gateway}}

Two items: OR - every pod in the edge namespaces, or app=gateway pods in the policy's own namespace. It passes review and functional tests, and lets far more in.

Read policies with k describe netpol: a ---------- separator between peers is the OR. And select "this exact namespace" with the automatic kubernetes.io/metadata.name label - custom namespace labels are only as trustworthy as whoever can set them.

Also asked: Design NetworkPolicies for a three-tier app behind an ingress controller. · Why does a policy with only ingress rules not restrict egress? · How would you debug "the NetworkPolicy is blocking us"?

Learn it: 16.31 Designing policies: three tiers, namespaces, the dash trap

What is the difference between the pod CIDR and the service CIDR, and how does a pod get its IP? Mid

Three ranges in every cluster, which must not overlap each other or anything the pods need to reach:

How a pod gets its IP: the kubelet asks containerd (CRI) for a pod sandbox (a network namespace held by the pause container); containerd runs the CNI plugin from /etc/cni/net.d; the plugin (Calico) creates a veth pair, takes an address from its IPAM, adds a node route to it; the kubelet writes status.podIP. Other nodes reach that pod block through a route (here IPIP via tunl0, learned over BGP).

If that fails: ContainerCreating and FailedCreatePodSandBox. Both ranges are hard to change later, so plan them first.

Also asked: What happens, network-wise, when a pod is created? · How does a pod on one node reach a pod on another node? · Why is the pod MTU smaller than the node's?

Learn it: 16.35 CNI: how a pod gets its IP, and why the CIDRs matter

Which volume would you use for scratch space, for files shared between containers, and for a database? Junior

A container's writable layer is thrown away on every container restart, so anything that matters goes in a volume (volumes: in the pod, volumeMounts: in the container):

Not hostPath: it ties the data to one node and, worse, a pod that can mount the node's filesystem can read the cluster's keys and other pods' secrets and take over the node. It is for node agents only.

Also asked: Why do security teams restrict hostPath? · A team's app loses its data every time the pod restarts. What do you check? · What is the difference between emptyDir and a PersistentVolumeClaim?

Learn it: 16.37 Volumes: where the data lives, and for how long

Explain PersistentVolume, PersistentVolumeClaim and StorageClass. Junior

Dynamic provisioning: create only the claim; the class's provisioner creates the PV (pvc-<uid>, rounded up in size). Static: an admin creates PVs for existing storage, and a claim binds when class, access modes, size, volume mode and selector all fit.

A claim stuck Pending: k describe pvc events - waiting for a pod (WaitForFirstConsumer, normal), no class or a broken provisioner, no matching static PV, or the default-class trap: a claim without storageClassName gets the default and never binds to a class-less PV - write storageClassName: "".

Also asked: A PVC is stuck in Pending. How do you figure out why? · What is the difference between static and dynamic provisioning? · What is a CSI driver?

Learn it: 16.39 PersistentVolume, PersistentVolumeClaim, StorageClass

What are the PersistentVolume access modes, and how do they affect scheduling? Junior

The backend decides what is possible: block storage (cloud disks, Ceph RBD) is RWO/RWOP only; shared file storage (NFS) can do RWX. So "three replicas share one volume" rules out disks.

Effects on scheduling:

Also asked: A Deployment with one replica and an RWO volume hangs during every rollout. Why? · What does WaitForFirstConsumer change? · What happens to a pod with a zonal disk when you drain the only node in that zone?

Learn it: 16.41 Access modes, zones, and why RWO constrains scheduling

What are the PersistentVolume reclaim policies, and what happens when someone deletes a PVC? Junior

The reclaim policy (persistentVolumeReclaimPolicy on the PV, copied from the StorageClass) decides what happens when the claim goes away:

Production data gets Retain: k patch pv NAME -p '{"spec":{"persistentVolumeReclaimPolicy":"Retain"}}' (a StorageClass's policy cannot be changed; make a new class).

To recover a Released volume: remove its stale claimRef (it holds the old claim's uid), so it becomes Available, then create a claim that binds to it by volumeName.

A PVC still used by a pod only goes Terminating - the pvc-protection finalizer waits for the pod. None of this is a backup: snapshots and tested restores are.

Also asked: Someone deleted a PVC holding production data. What can you do? · What happens to StatefulSet volumes when you scale down or delete it? · How do you grow a PVC, and can you shrink one?

Learn it: 16.44 Reclaim policies, protection, expansion, StatefulSet claims

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.