Kubernetes: Ingress, Gateway API & Service Mesh: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 34 of the course.
A user gets a 503 through your cluster's edge. How do you find which layer is failing? Mid
I follow the request hop by hop. First the edge: the ingress controller's or Gateway's log line for that request tells me whether a rule matched and which upstream it picked; a 503 from ingress-nginx usually means the matched Service has no ready endpoints, so I check kubectl get endpointslices and the pods' readiness. If there is a mesh, I read the response flag in the Envoy access log of the calling sidecar: UH means no healthy upstream, NC a missing subset, UO a circuit breaker, URX that retries were used up. The flag and the proxy that logged it tell me which hop and which object to fix, before I touch anything.
Also asked: What is the difference between Ingress and Gateway API? · How does cert-manager renew a certificate, and what do you check when it is stuck? · What does a service mesh give you, and what does it cost?
What is the difference between north-south and east-west traffic, and which tools handle each? Mid
North-south traffic enters or leaves the cluster: users calling shop.lab from outside. It goes through an edge proxy: a LoadBalancer Service at L4 and then an ingress controller or a Gateway at L7, which routes by host and path and usually terminates TLS. East-west traffic flows between Services inside the cluster; by default it is plain Service-to-Service networking, and a service mesh adds a proxy next to every pod for mTLS, retries, timeouts and uniform telemetry. When something fails, the first question is which of those hops answered.
Also asked: What is the difference between an L4 and an L7 load balancer? · Where can TLS terminate, and what changes with passthrough? · Which status codes would make you suspect the proxy rather than the application?
ingress-nginx is retired. What would you do about it in a company that depends on it? Mid
I would say what the retirement means first: since March 2026 the project ships no security fixes, so every new vulnerability stays open, but traffic keeps flowing. Then I would inventory: every Ingress, its annotations and the controller ConfigMap, because annotations are where the behaviour lives. I would run ingress2gateway to see what translates to Gateway API fields and what does not (snippets, some auth annotations), choose a Gateway API implementation, and move host by host, running both side by side behind the same DNS name so each cutover is a DNS change that can be rolled back.
Also asked: How do you see the nginx configuration ingress-nginx actually generated? · How does a canary Ingress choose which requests it gets? · Why are configuration snippets considered dangerous?
Learn it: 34.2 ingress-nginx in depth: the generated config, the ConfigMap, the annotations
Users get a 502 from ingress-nginx, but the pods look Running. What do you check? Mid
A 502 means nginx chose a pod and the connection broke, so I look at the controller's error log for that request first: connect() failed (111: Connection refused) means nothing listens on the targetPort, so I compare the Service's targetPort with the container port; recv() failed (104: Connection reset by peer) means the app closed the connection; an SSL handshake error suggests the backend speaks HTTPS and needs the backend-protocol annotation. I also check readiness, because Running is not Ready, and kubectl get endpointslices to see which pod addresses nginx can pick.
Also asked: What is the difference between a 503 and a 504 from an ingress controller? · How do you tell a 404 from the app from a 404 from the controller? · What causes an endless redirect loop behind a cloud load balancer?
Learn it: 34.3 Reading a 404, 502, 503 and 504 from ingress-nginx
A cert-manager Certificate has been not Ready for an hour. How do you debug it? Mid
I walk down the chain: Certificate, then the CertificateRequest it created, then the Order, then the Challenge, because the real error sits in the lowest object. cmctl status certificate <name> -n <ns> shows the whole chain. For an HTTP-01 Challenge the reason tells me what the self check saw: a DNS error means the name does not resolve to the ingress yet, a 404 means the solver Ingress is not picked up (often a wrong ingress class), a redirect means ssl-redirect is catching the challenge path. If the Issuer itself is not Ready, the ACME account could not register, for example because of the server URL or trust.
Also asked: What is the difference between HTTP-01 and DNS-01 challenges? · How does cert-manager decide when to renew a certificate? · What does the cert-manager.io/cluster-issuer annotation on an Ingress do?
Learn it: 34.10 Certificates that renew themselves: cert-manager and ACME
How does external-dns know which DNS records it is allowed to change? Mid
It uses a TXT registry: next to every record it creates, it writes a TXT record with its owner id, for example heritage=external-dns,external-dns/owner=<id> and the resource it came from. Before changing or deleting a record it checks that ownership record, so manually created names and records owned by another cluster are left alone. Together with the policy, upsert-only to never delete or sync to also clean up, that is what makes it safe to run against a shared zone; its log line, such as Adding RR, shows each change.
Also asked: What sources and providers does external-dns work with? · Why would a new Ingress get no DNS record? · How would you use DNS to cut traffic over from one ingress to another?
Learn it: 34.11 DNS that follows your Services: external-dns
An app team says its HTTPRoute is ignored. How do you find out why? Mid
I read the route's status first: kubectl describe httproute shows a condition per parent Gateway. Accepted False with NotAllowedByListeners means the listener's allowedRoutes does not admit the route's namespace; NoMatchingParent means the parentRef's name or sectionName does not exist; NoMatchingListenerHostname means the hostnames do not intersect. ResolvedRefs False with BackendNotFound means a missing Service or port, and RefNotPermitted means a cross-namespace backend without a ReferenceGrant. If the route is fine, I check the Gateway itself: Programmed True and an address. The status names the object to fix.
Also asked: What are the three roles in Gateway API, and which objects does each own? · Why does Gateway API need ReferenceGrant? · How is a Gateway listener different from an Ingress rule?
Learn it: 34.19 Gateway API in depth: roles, listeners, and what status tells you
How would you send 10% of traffic and all testers to a new version using Gateway API? Mid
One HTTPRoute with two rules. The first rule matches a header, for example x-canary: "on", and has a single backendRef to the v2 Service, so testers always reach v2. The second rule has no matches and two backendRefs, v1 with weight 90 and v2 with weight 10. Precedence picks the header rule for testers because it is more specific, whatever the order. I would add a timeouts.request so a slow v2 fails fast, watch v2's error rate, then move the weights step by step.
Also asked: How does Gateway API decide which rule matches a request? · What filters can an HTTPRoute rule apply? · What is the difference between rewriting a URL and redirecting it?
Learn it: 34.20 HTTPRoute traffic management: matches, filters, timeouts, mirrors
How would you migrate a cluster from Ingress to Gateway API without downtime? Mid
I would inventory the Ingresses and their annotations, then run ingress2gateway print to get Gateway and HTTPRoute YAML and read the warnings on stderr for what does not translate. After review I apply the Gateway next to the existing controller; it gets its own address, so I can test each host against it with a Host header or curl --resolve while users still hit the old edge. Then I move one host at a time by changing its DNS record, with a low TTL set in advance, watch error rates, and keep the old Ingress until the new path is proven. Rolling back is pointing DNS back.
Also asked: Which Ingress features have no direct Gateway API equivalent? · How do you test a new ingress path before users reach it? · What does ingress2gateway produce, and what does it not do?
What does a service mesh give you, and what does it cost? Mid
A mesh puts a proxy next to every workload, the data plane, and a control plane configures those proxies and issues each workload an identity certificate. Without changing the apps I get mTLS and identity-based authorization between services, consistent retries, timeouts and traffic splitting, and the same request metrics for every service. The costs are real: memory and CPU per pod, extra latency per hop, a complex system with its own upgrades and failure modes, and debugging that now includes the proxies. I would recommend one when there are many services and a security or reliability requirement it solves, not by default.
Also asked: What is the difference between the data plane and the control plane? · What is xDS, and what does the control plane send over it? · When would you decide against a service mesh?
Learn it: 34.27 What a service mesh gives you, and what it costs
You labelled a namespace for Istio injection but traffic is not going through the mesh. Why? Mid
Injection happens when a pod is created: a mutating webhook adds the proxy to new pods in a namespace labelled istio-injection=enabled. Pods that already ran stay 1/1, so I check kubectl get pods for 2/2 and run istioctl analyze, which reports IST0103 for pods without a sidecar. The fix is kubectl rollout restart on the Deployments. Then istioctl proxy-status should list every new pod as synced with istiod. If the restarted pods fail to be created at all, the webhook is failing closed, usually because istiod is down.
Also asked: How does Istio redirect a pod's traffic through the sidecar? · What happens to new pods if istiod is unavailable? · What does istioctl analyze check?
Learn it: 34.28 Istio: install, injection and the control plane
You turned on STRICT mTLS and some clients broke. How do you work out which and fix it? Mid
PeerAuthentication in STRICT means the workload accepts only mTLS, so clients without a sidecar now get a connection reset, curl's (56) Recv failure, and the server's proxy logs NR filter_chain_not_found. Those log lines give me the source addresses of the plain-text clients. The safe path is back to PERMISSIVE, then either put those clients in the mesh or, if one port must stay open, use portLevelMtls for that port only, then STRICT again. If a client gets 403 RBAC: access denied instead, mTLS works and an AuthorizationPolicy denies it.
Also asked: What is a SPIFFE ID, and where does it come from? · How does Istio evaluate ALLOW and DENY AuthorizationPolicies? · What is the difference between PERMISSIVE and STRICT mode?
What is a retry storm, and how do you prevent one in a service mesh? Mid
When every layer of a call chain retries, one failing service gets the product of all those retries: with Istio's default of 2 attempts at three layers, a request can become up to 27 calls to the bottom service, which makes it fail harder. To prevent it I retry at one layer, usually the one closest to the failing dependency, and set retries.attempts: 0 in the VirtualServices of the other layers. Timeouts get shorter the deeper the call goes, so outer layers do not give up while inner ones still work. Outlier detection and connection pool limits stop sending traffic to endpoints that are already failing.
Also asked: What is the difference between a VirtualService and a DestinationRule? · How would you test a timeout with fault injection? · What does outlier detection do?
Learn it: 34.35 Traffic policy: VirtualService, DestinationRule, retries, timeouts, faults
A call between two services fails with a 503 in the mesh. How do you find the cause? Mid
I read the access log of the client's sidecar with kubectl logs <pod> -c istio-proxy and look at the response flag. UH means no healthy upstream: the subset or Service has no ready endpoints, which istioctl proxy-config endpoints confirms. NC means no cluster: the route points at a subset no DestinationRule defines. UO means a connection pool limit tripped, URX that retries were used up on a failing backend, UF that it could not connect at all. Then I check the server's inbound line: if there is none, the request never arrived. The flag tells me which object to fix.
Also asked: Why does one request produce two access log lines in a mesh? · What does istioctl proxy-config show you? · How would you see whether a proxy is retrying requests?
Learn it: 34.36 Reading Envoy: the access log, response flags and proxy-config
What telemetry does a service mesh give you without changing the applications? Mid
Every proxy reports the same request metrics for every service: a request count with labels for source, destination and response code, and request duration, so I get rate, errors and latency per service and per caller in any language. Access logs per request come from the proxies too, and the Telemetry API can turn them on per namespace. What the mesh cannot do alone is tracing across hops: the app must copy the trace headers onto its outgoing calls. To check latency myself I run fortio and read the target 99% line, in seconds, because the tail moves before the median.
Also asked: What labels do the standard Istio request metrics carry? · Why do p99 and the median tell different stories? · Why must applications propagate trace headers?
Learn it: 34.37 Mesh telemetry: golden signals from every proxy
What is Istio ambient mode, and how does it differ from sidecars? Mid
In ambient mode there is no proxy in the pod. A ztunnel on each node handles L4 for all ambient pods on it: mTLS, workload identity and L4 authorization. L7 features (HTTP routing, retries, header-based policies) come from an optional waypoint proxy per namespace or Service. Compared with sidecars that means much less memory, no pod restarts to join or upgrade, and paying for L7 only where it is used; the costs are a node-wide blast radius for ztunnel and an extra hop through the waypoint for L7 traffic. Both modes can run in one mesh, so migration can be gradual.
Also asked: How does Linkerd differ from Istio? · What is GAMMA in Gateway API? · How would you decide whether a team needs a service mesh at all?
Learn it: 34.38 Ambient mode, Linkerd, and when not to use a mesh
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.