In plain words
Imagine your toolbox is in the garage, but the thing you need to fix is inside a sealed glass box. You can only reach in through a small hatch. If someone packed a screwdriver inside the box, you can use it through the hatch. If they packed nothing, you need to slide a second little box with tools right next to it, sharing the same air.
In Kubernetes the JVM runs in the container as PID 1, and kubectl exec POD -- jcmd 1 Thread.print is reaching through the hatch. If the image has no JDK, kubectl debug --target=app adds an ephemeral container with tools that shares the pod's process namespace, so it can see PID 1, as long as it runs as the same user.
PID 1 is the JVM
The problem. Everything in this chapter so far ran on the VM. In a pod the JVM is PID 1 of a container whose image may not even contain the tools - and you still have to get a thread dump out of it at 3am.
What you need to know already: the JDK tools (20.3), kubectl exec, kubectl cp and debugging (15.42), requests/limits and probes for JVMs (17.6, 17.22), PID 1 in containers (10.28).
In a container with an exec-form entrypoint (chapter 10), the JVM is PID 1:
# in a cluster, on a JDK-based image (pod name is an example)
kubectl exec -it orders-7d9f6c5b8-2xkpq -- jcmd
1 /app/orders.jar
78 jdk.jcmd/sun.tools.jcmd.JCmd
kubectl exec orders-7d9f6c5b8-2xkpq -- jcmd 1 GC.heap_info
kubectl exec orders-7d9f6c5b8-2xkpq -- jcmd 1 Thread.print > dump1.txt
Everything from this chapter works the same, with kubectl exec ... -- in front and PID 1. It works because the exec runs as the same user as the JVM, in the same /tmp.
When the image has no tools
kubectl exec orders-... -- jcmd 1 Thread.print
error: Internal error occurred: ... exec: "jcmd": executable file not found in $PATH
JRE and distroless images. Options, from quickest to most prepared:
kill -3 inside the pod: kubectl exec POD -- kill -3 1 (needs a kill binary; distroless has none)
Actuator: kubectl port-forward POD 8080 & curl -H 'Accept: text/plain' localhost:8080/actuator/threaddump
Ephemeral debug container: kubectl debug -it POD --image=eclipse-temurin:21-jdk --target=app -- jcmd 1 Thread.print
Ship a JDK runtime image: bigger image, every tool always there
kubectl debug --target=<container> puts the debug container in the target's process namespace, so it sees the JVM as PID 1. The attach mechanism still needs the same UID and access to the JVM's /tmp (through /proc/1/root/tmp): run the debug container as the same user (for example with a --profile or an image whose default user matches) or you get the "Operation not permitted" you saw on the VM.
Getting files out
kubectl exec POD -- jcmd 1 GC.heap_dump /tmp/heap.hprof
kubectl cp NAMESPACE/POD:/tmp/heap.hprof ./heap.hprof
Mind the size: /tmp in a container is usually the writable layer, counted against ephemeral-storage and sometimes the memory limit (if it is a tmpfs emptyDir: {medium: Memory}). A heap dump to a memory-backed tmpfs can get the pod OOM-killed while you are dumping it. Write dumps to a real volume.
The flags, as a Deployment
env:
- name: JAVA_TOOL_OPTIONS
value: >-
-XX:MaxRAMPercentage=75
-XX:+ExitOnOutOfMemoryError
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/dumps
-Xlog:gc*:file=/dumps/gc.log:time,uptime:filecount=5,filesize=10M
resources:
requests: { memory: 1Gi, cpu: 500m }
limits: { memory: 1Gi }
volumeMounts:
- { name: dumps, mountPath: /dumps }
ExitOnOutOfMemoryError matters more in Kubernetes than anywhere: a JVM that survives an OOM with half its threads dead still answers the liveness probe and never gets restarted.
CPU limits and the JVM
The JVM sizes GC and compiler threads, and the common pool, from the CPU count it sees - which honours CPU limits (quota), not requests. A pod with limits.cpu: 1 gets availableProcessors() == 1: Serial GC (below 2 CPUs), ForkJoinPool.commonPool parallelism 1 (CPUs - 1, but never below 1), and CompletableFuture falling back to a thread per task, as it does whenever the common pool's parallelism is 1. -XX:ActiveProcessorCount=2 overrides it when you know better.
FAQ
Why is the JVM PID 1 in the container?
Because the image's entrypoint is in exec form, like ["java", "-jar", "/app/orders.jar"], so the JVM is the first process started in the container's PID namespace. With a shell-form entrypoint, sh -c would be PID 1 and the JVM its child, and signals like SIGTERM might not reach it. For the tools this means jcmd 1 targets the JVM, and kubectl exec runs as the same user in the same /tmp, so attach works.
How do I get a thread dump if the image has no JDK?
Several ways. If the image has a shell and kill, kubectl exec POD -- kill -3 1 prints the dump to the container's stdout, which you read with kubectl logs. If Spring Boot Actuator is exposed on a management port, kubectl port-forward plus curl localhost:8080/actuator/threaddump. Otherwise an ephemeral debug container with a JDK: kubectl debug -it POD --image=eclipse-temurin:21-jdk --target=app -- jcmd 1 Thread.print.
Why does jcmd in a debug container say "Operation not permitted"?
The attach mechanism needs the same user as the JVM and access to its /tmp. A debug container often runs as root or a different UID than the application container. Root may still fail when the target's security context blocks it or capabilities are dropped. Run the debug container as the application's UID, through a debug profile or an image whose default user matches, and the tool reaches the socket through /proc/1/root/tmp.
Can a heap dump get my pod killed?
Yes. If /tmp or the dump directory is a memory-backed emptyDir (medium: Memory), the file counts against the pod's memory limit, and a heap-sized file on top of a heap-sized process exceeds it: OOMKilled mid-dump. Even on disk, the writable layer counts against ephemeral-storage limits and can trigger eviction. Write dumps to a proper volume and copy them out with kubectl cp before the pod goes away.
Do CPU requests or limits decide what the JVM sees?
Limits. The JVM reads the CPU quota from the cgroup, which comes from limits.cpu, and treats it as the processor count. Requests only affect scheduling and CPU shares, not what availableProcessors() returns. With limits.cpu: 1 you get one processor: Serial GC, fewer GC and compiler threads, and a ForkJoin common pool of parallelism 1. -XX:ActiveProcessorCount=2 overrides that. Without a CPU limit, the JVM sees the node's CPUs.
In an interview Mid
How do you troubleshoot a Java application running in a Kubernetes pod?
The same tools, with kubectl exec POD -- in front and PID 1 (the JVM is the entrypoint): kubectl exec POD -- jcmd 1 Thread.print > dump1.txt, jcmd 1 GC.heap_info, jcmd 1 VM.flags. It works because exec runs as the JVM's user, in the same /tmp.
When the image has no tools (JRE or distroless, executable file not found):
kubectl exec POD -- kill -3 1 and read kubectl logs (needs a kill binary).- An ephemeral container:
kubectl debug -it POD --image=eclipse-temurin:21-jdk --target=app - it shares the process namespace; it must run as the same UID or attach fails.
Files out: kubectl cp NAMESPACE/POD:/tmp/heap.hprof . - write dumps to a real volume, not a memory-backed emptyDir.
And check the flags: JAVA_TOOL_OPTIONS with MaxRAMPercentage=75, ExitOnOutOfMemoryError (or a half-dead JVM keeps passing liveness), and CPU limits: limits.cpu: 1 means Serial GC and a one-thread common pool.
Also asked: What happens to a JVM when its pod exceeds the memory limit? · Why does ExitOnOutOfMemoryError matter more in Kubernetes than on a VM? · How do CPU limits change what the JVM decides at startup?