OnCallReady

Lesson 15.1 · Kubernetes: Architecture & Workloads · 25 min read

The lab cluster, kubectl and kubeconfig

In plain words

Think of a TV remote. It doesn't control the TV by magic; it needs to know which TV to point at, and it needs to be paired with it. If you have several TVs in the house, the remote has a switch for "living room" or "bedroom". Point it at nothing and pressing buttons does nothing.

kubectl is the remote, and the kubeconfig file (~/.kube/config) is the pairing: it lists clusters (where: the apiserver URL and its CA), users (who: here a client certificate that is cluster-admin) and contexts (a named cluster-plus-user pair), with current-context as the switch. kubectl talks only to the kube-apiserver at https://10.64.0.10:6443. The famous localhost:8080 refused error means kubectl has no kubeconfig at all, not that the cluster is down.

Why Kubernetes

In Ch 10-11 you ran containers with docker run and Docker Compose - on one machine. If that machine dies, your containers die with it. If you need 20 copies of the API spread over 5 machines, or a new version rolled out without downtime, Docker alone gives you no help: you would be ssh-ing into machines and running docker run by hand.

Kubernetes (often written K8s: K, eight letters, s) solves that. You give it a list of machines and a description of what should run ("3 copies of orders:2.4, always"), and it decides which machine runs what, starts the containers, restarts them when they crash and replaces them when a machine dies. It is Terraform's idea (Ch 12: declare what you want, a tool makes reality match) applied to running containers, and it never stops checking.

What you need to know already: containers and images (10.3, 10.5), docker run / logs / exec (11.1, 11.5), systemd services (2.1), HTTPS, certificates and CAs (9.15), environment variables and sudo (1.11).

The words for the pieces

What you are connecting to

From this chapter on, oncall-lab is your workstation - you type commands here. The thing you operate is a three-node cluster, built with kubeadm (the official tool that installs Kubernetes on plain Linux machines):

cp-1       10.64.0.10   control plane
worker-1   10.64.0.11   worker
worker-2   10.64.0.12   worker

Some facts about it you will meet again:

The cluster is pre-built so you can start with workloads.

Later (Ch 18): you build a cluster like this yourself with kubeadm init and kubeadm join.

The one rule: kubectl talks only to the API server

The control plane has a program called kube-apiserver (the API server, or apiserver). It is an HTTPS API (9.21): every piece of the cluster is a record you can read or change through it. Keep one fact in your head from minute one:

kubectl talks to exactly one thing: the kube-apiserver. Never to a node, never to the database, never to a container runtime.

kubectl get pods (list the pods) is just an HTTPS GET to https://10.64.0.10:6443/api/v1/namespaces/default/pods - 6443 is the API server's port. kubectl apply (create or update something from a file) is a POST or PATCH. The API server stores everything in etcd, a small database on the control plane (15.5 explains it).

Everything that then happens on the nodes happens because some other program watches the API server and reacts. That is the whole architecture; the rest of this chapter makes you watch it happen.

How kubectl knows where the cluster is: kubeconfig

kubectl needs two things: where the API server is and who you are. It reads them from a kubeconfig file. It looks, in this order, at:

  1. --kubeconfig=PATH on the command line
  2. the KUBECONFIG environment variable (several files separated by :, merged)
  3. ~/.kube/config - the normal place

With none of them, it falls back to http://localhost:8080 - a default from very old versions - and you get the most common first error in Kubernetes:

$ kubectl get pods
E0923 10:14:02.118431    4012 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"http://localhost:8080/api?timeout=32s\": dial tcp 127.0.0.1:8080: connect: connection refused"
E0923 10:14:02.119907    4012 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: ...
...
The connection to the server localhost:8080 was refused - did you specify the right host or port?

Read it correctly: localhost:8080 means "I have no kubeconfig", not "the cluster is down". The repeated E0923 lines (E = error, 0923 = the date, Sept 23) are kubectl retrying. "Connection refused" is the TCP failure from 9.1: nothing listens on port 8080 of your own machine.

If the message names the real server (10.64.0.10:6443) instead, then the API server is unreachable - a completely different problem.

Two variations of the same mistake you will make at some point:

$ sudo kubectl get pods
The connection to the server localhost:8080 was refused - did you specify the right host or port?

sudo runs kubectl as root, whose home directory is /root - so it looks for /root/.kube/config, which does not exist. You never need sudo for kubectl: your permissions come from the credentials inside the kubeconfig, which the API server checks - not from your Linux user.

$ KUBECONFIG=/tmp/nothere kubectl get pods
The connection to the server localhost:8080 was refused - did you specify the right host or port?

An environment variable pointing at a missing file behaves the same way.

Anatomy of a kubeconfig

When kubeadm installs a cluster it writes an admin kubeconfig to /etc/kubernetes/admin.conf on the control plane node, and prints these three commands to copy it for your user:

mkdir -p $HOME/.kube
sudo cp -i /etc/kubernetes/admin.conf $HOME/.kube/config
sudo chown $(id -u):$(id -g) $HOME/.kube/config

(mkdir -p makes the directory; cp -i asks before overwriting; chown $(id -u):$(id -g) gives the copy to your user and group, 4.3.)

The file is YAML - the indented key: value format you wrote for Compose in 11.31. A line starting with - is a list item. It has three lists and one pointer:

apiVersion: v1
kind: Config
clusters:                       # WHERE: server URL + the CA that signed its cert
- cluster:
    certificate-authority-data: LS0tLS1CRUdJTi...
    server: https://10.64.0.10:6443
  name: kubernetes
users:                          # WHO: credentials (here a client certificate)
- name: kubernetes-admin
  user:
    client-certificate-data: LS0tLS1CRUdJTi...
    client-key-data: LS0tLS1CRUdJTi...
contexts:                       # a named pair: cluster + user (+ default namespace)
- context:
    cluster: kubernetes
    user: kubernetes-admin
  name: kubernetes-admin@kubernetes
current-context: kubernetes-admin@kubernetes     # which context is active

That client-key-data is a private key that grants cluster-admin: permission to do anything to anything in the cluster. Treat the file like an SSH private key (1.17): chmod 600 (only you can read it, 4.5), never commit it, never paste it in a ticket. kubectl does not warn about a world-readable kubeconfig, so nothing stops you except habit.

Reading and switching contexts

kubectl config is the group of subcommands that read and change the kubeconfig. They never talk to the cluster.

# after the next mission has installed ~/.kube/config
kubectl config get-contexts
CURRENT   NAME                          CLUSTER      AUTHINFO           NAMESPACE
*         kubernetes-admin@kubernetes   kubernetes   kubernetes-admin

get-contexts lists every context. CURRENT has a * on the active one; AUTHINFO is the user entry; NAMESPACE is empty, meaning "default".

kubectl config current-context
kubernetes-admin@kubernetes

current-context prints just the active one's name.

kubectl config view
apiVersion: v1
clusters:
- cluster:
    certificate-authority-data: DATA+OMITTED
    server: https://10.64.0.10:6443
  name: kubernetes
...

config view prints the merged kubeconfig with the secrets hidden (DATA+OMITTED). --raw shows them. --minify shows only what the current context uses - handy when a merged KUBECONFIG has ten clusters in it. kubectl config use-context NAME switches the active context.

At work you will usually have one context per cluster (dev, test, acc, prod). The single most expensive kubectl mistake is running a command against the wrong one. Run kubectl config current-context before anything destructive. In the Kubernetes admin certification exam in your plan (a hands-on test, not a quiz), every task starts with a kubectl config use-context ... line you must run.

First contact

# once ~/.kube/config is in place (the next mission)
kubectl cluster-info
Kubernetes control plane is running at https://10.64.0.10:6443
CoreDNS is running at https://10.64.0.10:6443/api/v1/namespaces/kube-system/services/kube-dns:dns/proxy

To further debug and diagnose cluster problems, use 'kubectl cluster-info dump'.

cluster-info proves you reach the API server and prints its URL. CoreDNS is the cluster's own DNS server (8.16) - it lets pods find each other by name.

kubectl get nodes -o wide
NAME       STATUS   ROLES           AGE   VERSION   INTERNAL-IP     EXTERNAL-IP   OS-IMAGE             KERNEL-VERSION     CONTAINER-RUNTIME
cp-1       Ready    control-plane   12d   v1.34.1   10.64.0.10   <none>        Ubuntu 24.04.3 LTS   6.8.0-85-generic   containerd://2.1.4
worker-1   Ready    <none>          12d   v1.34.1   10.64.0.11   <none>        Ubuntu 24.04.3 LTS   6.8.0-85-generic   containerd://2.1.4
worker-2   Ready    <none>          12d   v1.34.1   10.64.0.12   <none>        Ubuntu 24.04.3 LTS   6.8.0-85-generic   containerd://2.1.4

kubectl get nodes lists the nodes; -o wide ("output wide") adds extra columns. Column by column:

columnmeans
NAMEthe node's name
STATUSReady = the node's agent reports it can run pods
ROLEScontrol-plane for cp-1; workers show <none> (kubeadm only labels the control plane)
AGEhow long ago the node joined (12 days)
VERSIONthe Kubernetes version of the agent on that node, not of the cluster
INTERNAL-IPthe node's IP on the cluster network
EXTERNAL-IPa public IP, if the machine had one
OS-IMAGE, KERNEL-VERSIONwhat cat /etc/os-release and uname -r would say on the node
CONTAINER-RUNTIMEthe runtime that starts containers there: containerd 2.1.4

The agent on each node is called the kubelet (15.7). It writes this information into the API; you are reading the stored copy.

kubectl version
Client Version: v1.34.1
Kustomize Version: v5.7.1
Server Version: v1.34.1

kubectl version shows the client (your kubectl) and the server (the API server) versions. They are separate programs. Kubernetes versions are major.minor.patch; kubectl supports one minor version of difference ("skew") either way - a 1.34 kubectl works with 1.33, 1.34 and 1.35 servers. Outside that window things mostly work until they subtly do not. (Kustomize is a YAML tool built into kubectl; ignore it for now.)

The oncall-lab kubelet is a different thing

In the node-NotReady incident (2.36) a kubelet.service ran on oncall-lab itself. That is a node agent that was never joined to this cluster - it has nothing to do with kubectl. Don't let the names confuse you: kubelet = the agent that runs on each node; kubectl = the client you type.

What you can now do:

Why it helps

The most expensive kubectl mistake at work is running a command against the wrong cluster, prod instead of test. Knowing contexts and checking kubectl config current-context before anything destructive is the habit that prevents it, and every task in the exam (Ch 19) starts with a use-context line you must run. You'll also hit the classic confusions: sudo kubectl failing because root has no kubeconfig, localhost:8080 meaning "no config" rather than "cluster down", and a kubeconfig containing a cluster-admin private key that must be treated like an SSH key. Reading kubectl get nodes -o wide correctly (kubelet version per node, runtime) is the first thing you'll do on any new cluster.

Commands in this lesson

kubectl

FAQ

Why does kubectl say localhost:8080 connection refused?

kubectl found no kubeconfig: no --kubeconfig, no valid KUBECONFIG, no ~/.kube/config. It falls back to an unauthenticated http://localhost:8080 from very old clusters. So the message means "I don't know where the cluster is", not "the cluster is down". If the error names the real server, like 10.64.0.10:6443, then the apiserver is genuinely unreachable.

Why does sudo kubectl fail when kubectl works?

sudo runs kubectl as root, whose HOME is /root, so it looks for /root/.kube/config, which doesn't exist. You never need sudo for kubectl: your permissions come from the credentials in the kubeconfig, checked by the apiserver's RBAC, not from your Unix user on the workstation.

Is the kubeconfig file sensitive?

Very. The kubeadm admin kubeconfig contains a client key that grants cluster-admin: full control of every object, including secrets. Treat it like an SSH private key: chmod 600, never commit it, never paste it into a ticket. kubectl config view redacts the data (DATA+OMITTED); --raw shows it.

What does the VERSION column in kubectl get nodes mean?

It's the kubelet version on that node, not the cluster's version. The control plane version is the Server Version from kubectl version. During upgrades they differ, and Kubernetes allows kubelets to lag the apiserver by a limited number of minor versions. kubectl itself supports one minor version of skew either way.

How do I work with several clusters safely?

Keep one context per cluster, check kubectl config current-context before any destructive command, and make the context visible in your prompt. KUBECONFIG can list several files separated by colons, merged; kubectl config view --minify shows only what the current context uses. kubectl config use-context NAME switches.

In an interview Junior

What is a kubeconfig file, and how does kubectl find it?

A kubeconfig is the YAML file that tells kubectl where the cluster is and who you are. It has three lists and a pointer:

kubectl looks at --kubeconfig=PATH, then the KUBECONFIG variable (several files merged), then ~/.kube/config. With none it falls back to localhost:8080 - so "The connection to the server localhost:8080 was refused" means no kubeconfig, not a down cluster (and sudo kubectl looks in /root).

kubectl config get-contexts, current-context and use-context read and switch; check the context before anything destructive.

Also asked: kubectl says "The connection to the server localhost:8080 was refused". What does that mean? · How do you avoid running a command against the wrong cluster? · What does kubectl version tell you, and what is version skew?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.