Why cluster operations look different here
On the kubeadm cluster you upgraded node by node and SSHed into machines (18.16, 18.1). On OpenShift you ask for a version and the cluster does the rest; the nodes have no SSH for you; and Red Hat support wants a specific bundle of data. Three new habits.
What you need to know already: kubeadm upgrades and version skew (18.15, 18.16); drain, cordon and PodDisruptionBudgets (18.25, 17.26); ClusterVersion and ClusterOperators (30.1); the kubelet, CRI and crictl (15.7, 18.1); journalctl -u and systemctl is-active (2.x systemd chapter); etcd backups (18.20).
The words you need first
- Release image - one image whose content is the list of exact images of every platform component for one OpenShift version.
- Update channel - which stream of releases the cluster is offered (
stable-4.21). - z-stream - a patch upgrade within a minor (4.21.25 -> 4.21.27); a minor upgrade changes the middle number (4.21 -> 4.22).
- MachineConfig / MachineConfigPool - OS configuration for nodes, and the group of nodes (masters, workers) it applies to. The Machine Config Operator rolls them out.
- must-gather - a command that collects the cluster's state and logs into one directory for support.
The cluster upgrades itself
An OpenShift release is a release image (quay.io/openshift-release-dev/ocp-release@sha256:...) that lists the exact image of every platform component. The Cluster Version Operator (CVO) applies it: it rolls out new manifests for each cluster operator in order, and the operators update their operands. Last, the Machine Config Operator updates the nodes' OS (RHCOS) and kubelet, draining and rebooting nodes one at a time per MachineConfigPool. You do not upgrade etcd, the apiserver or the kernel by hand; you ask for a version.
Where you are
oc adm upgrade with no arguments only reads: current version, channel, and the updates on offer:
$ oc get clusterversion
NAME VERSION AVAILABLE PROGRESSING SINCE STATUS
version 4.21.25 True False 12d Cluster version is 4.21.25
$ oc adm upgrade
Cluster version is 4.21.25
Upstream is unset, so the cluster will use an appropriate default.
Channel: stable-4.21 (available channels: candidate-4.21, candidate-4.22, eus-4.22, fast-4.21, fast-4.22, stable-4.21, stable-4.22)
Recommended updates:
VERSION IMAGE
4.21.27 quay.io/openshift-release-dev/ocp-release@sha256:b466...
4.21.26 quay.io/openshift-release-dev/ocp-release@sha256:ae64...
The CVO asks Red Hat's update service (the "upstream", api.openshift.com, or a local OpenShift Update Service in disconnected clusters) which releases are recommended from the current version on the current channel. The update graph knows which paths are tested and which have known issues (those appear as "conditional updates" with a risk explanation, not in the recommended list).
Channels
For each minor there are:
- candidate-4.x - every release as soon as it is built; for testing, not support.
- fast-4.x - releases once Red Hat declares them GA; fully supported.
- stable-4.x - the same releases, promoted after they have soaked on fast clusters (telemetry shows no problems). What production normally uses.
- eus-4.x - even minors only (4.18, 4.20, 4.22): long-support releases, and the basis for EUS-to-EUS upgrades (e.g. 4.20 -> 4.22 with only the control plane passing through 4.21, worker pools paused to avoid a second reboot).
Moving to the next minor is two steps: switch to that minor's channel, then pick an update the graph offers:
$ oc adm upgrade channel stable-4.22
$ oc adm upgrade
...
Recommended updates:
VERSION IMAGE
4.22.4 quay.io/openshift-release-dev/ocp-release@sha256:...
4.21.27 quay.io/openshift-release-dev/ocp-release@sha256:...
(simulator) The lab's update graph is lab data; the real one is on https://access.redhat.com/labs/ocpupgradegraph/. Before a minor upgrade you also check the release notes for removed Kubernetes APIs (the cluster sometimes requires an admin-ack - oc adm upgrade then shows Upgradeable=False with the reason until a cluster admin acknowledges it in the admin-acks ConfigMap) and that every add-on operator supports the target version.
Running one
$ oc adm upgrade --to-latest
Requested update to 4.21.27
$ oc get clusterversion
NAME VERSION AVAILABLE PROGRESSING SINCE STATUS
version 4.21.25 True True 20s Working towards 4.21.27: 252 of 889 done (28% complete), waiting on config-operator
$ oc get co | grep -v '4.21.25'
NAME VERSION AVAILABLE PROGRESSING DEGRADED SINCE MESSAGE
etcd 4.21.27 True False False 13d
kube-apiserver 4.21.27 True False False 13d
...
--to=4.21.27 picks a version; --to-latest the newest recommended. The status counts the release's manifests; VERSION stays the old one until the upgrade completes. During the upgrade: operators flip to the new version one by one; a cluster operator that stays Progressing or turns Degraded is where you look (oc describe co <name> - its conditions carry a message). The last and longest phase is machine-config: node by node, drain (PDBs matter - 17.26), reboot, uncordon. (grep -v '4.21.25' above hides the lines still at the old version, so you see only the operators that already moved.) oc get nodes shows SchedulingDisabled on the node being updated; oc get mcp (MachineConfigPools) shows UPDATED/UPDATING per pool. (simulator) The lab has no MachineConfigPools; its "upgrade" finishes when the operators do.
When it is done:
# on an idle cluster (the upgrade mission starts one)
oc adm upgrade
Cluster version is 4.21.27
...
oc get clusterversion version -o jsonpath='{range .status.history[*]}{.version}{" "}{.state}{" "}{.completionTime}{"\n"}{end}'
4.21.27 Completed 2026-09-24T10:14:51Z
4.21.25 Completed 2026-09-11T10:50:31Z
4.21.22 Completed 2026-08-12T08:50:31Z
Upgrades go forward only. There is no downgrade of an OpenShift cluster; the safety net is testing on a non-prod cluster in the same channel first, and etcd backups for disasters.
Nodes: no SSH, oc debug node
RHCOS is an immutable, image-based OS (rpm-ostree: the whole OS is updated as one image, not package by package): the Machine Config Operator owns its files. SSH to nodes is possible (the core user with an installer key) but discouraged - changes made by hand are either reverted by MCO or mark the node as "tainted" for support. The supported way in is a debug pod:
$ oc debug node/worker-1
Temporary namespace openshift-debug-fwjhg is created for debugging node...
Starting pod/worker-1-debug-7nxgk ...
To use host binaries, run `chroot /host`. Instead, if you need to access host namespaces, run `nsenter -a -t 1`.
Pod IP: 10.64.0.11
If you don't see a command prompt, try pressing enter.
sh-5.1# crictl ps
sh: crictl: command not found
sh-5.1# chroot /host
sh-5.1# crictl ps --name router
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD NAMESPACE
b2866e113b69d sha256:a7b120f752e8b... 13 days ago Running router 0 jjk9r87svkk9z router-default-7884t2w448-bmbmh openshift-ingress
sh-5.1# journalctl -u kubelet -n 5
sh-5.1# systemctl is-active kubelet crio
sh-5.1# exit
sh-5.1# exit
Removing debug pod ...
Temporary namespace openshift-debug-fwjhg was removed.
What happened: a privileged pod on that node (in a temporary namespace, host network, host PID) with the node's root filesystem mounted at /host. The debug image (a RHEL support-tools image) has its own tools; chroot /host ("change root": make /host the / of this shell) switches to the node's binaries - crictl, journalctl, systemctl, rpm-ostree. The first exit leaves the chroot, the second leaves the pod. One-shot form for scripts (everything after -- runs instead of a shell):
$ oc debug node/worker-1 -- chroot /host journalctl -u kubelet -n 20 --no-pager
This needs cluster-admin (it creates a privileged pod); developers get Forbidden. (simulator) The lab nodes are the kubeadm lab cluster's nodes, presented as RHCOS inside the debug shell; a handful of host commands are simulated there.
must-gather
When you open a Red Hat support case, the first thing they ask for:
$ oc adm must-gather
[must-gather ] OUT 2026-09-24T10:12:03.123456789Z Using must-gather plug-in image: quay.io/openshift-release-dev/ocp-v4.0-art-dev@sha256:...
When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information:
ClusterID: 6f2d7c1e-8a41-4b7d-9e3a-2c5f0b9d4e17
ClientVersion: 4.21.25
ClusterVersion: Stable at "4.21.25"
ClusterOperators:
All healthy and stable
[must-gather ] OUT 2026-09-24T10:12:03.2Z namespace/openshift-must-gather-8fh4x created
[must-gather ] OUT 2026-09-24T10:12:03.3Z clusterrolebinding.rbac.authorization.k8s.io/must-gather-v9d5k created
[must-gather ] OUT 2026-09-24T10:12:03.4Z pod for plug-in image ... created
[must-gather-5xdjq] POD 2026-09-24T10:12:04Z Gathering data for ns/openshift-cluster-version...
...
[must-gather ] OUT 2026-09-24T10:12:40Z namespace/openshift-must-gather-8fh4x deleted
Reprinting Cluster State:
...
$ ls must-gather.local.*/
quay-io-openshift-release-dev-ocp-v4-0-art-dev-sha256-...
It runs a gather pod that collects cluster-scoped resources, every openshift-* namespace's objects and pod logs, and node data, then copies it to must-gather.local.<number>/ (--dest-dir to choose), which you tar and upload. Product-specific images add more (--image=registry.redhat.io/... for logging, OLM, storage...). oc adm inspect ns/shop is the narrow version: one namespace or resource, same layout. Both only read; both need cluster rights. The output contains Secrets' names and configs but not their values - still treat it as sensitive.
What you can now do
- Read the version, channel and recommended updates, run a z-stream upgrade and follow it through
oc get clusterversionandoc get co. - Get a root shell on an RHCOS node with
oc debug node+chroot /host. - Collect
oc adm must-gatherandoc adm inspectfor a support case.