Imagine timing how long it takes a family to get out of the house in the morning. A list of how long each person spent in the bathroom is interesting, but people use different bathrooms at the same time. What actually made you late is the chain: you waited for your sister, who waited for the hairdryer, which waited for dad.
systemd-analyze blame is the bathroom list: how long each unit took to start. systemd-analyze critical-chain is the chain of waiting that really delayed the boot, often ending at a unit that waits for the network. The other tools are checks and explanations: verify proofreads a unit file, security scores its sandbox, calendar explains a timer schedule. daemon-reexec is systemd restarting itself without a reboot.
Why this matters
"The server takes two minutes to come back after a reboot." Is it the hardware, the kernel, or one slow service everybody waits for? systemd-analyze measures the boot and shows you which chain of units actually held it up - which is usually not the unit that took longest.
What you need to know already: 2.14 (ordering with After=, targets), 2.7 (systemd-analyze verify).
Where did the boot go
The boot has two halves: the kernel part (the kernel starting up and finding the hardware) and the userspace part (everything after the kernel hands over to systemd: all the units).
$systemd-analyze time
Startup finished in 1.842s (kernel) + 4.127s (userspace) = 5.969s
graphical.target reached after 4.066s in userspace.
blame lists each unit with how long it took to start, slowest first. (On this lab box, kubelet and orders are two programs installed as services; you meet kubelet in 2.36. apparmor loads Ubuntu's security rules. | head -4 keeps only the first 4 lines.)
blame is not the same as slow boot. Units start in parallel - a 2-second service that nothing waits for costs you nothing.
Read it bottom-up: each unit had to wait for the one below it. @ is when the unit finished starting (seconds after userspace began); + is how long it took. That is the chain that actually delayed you, and systemd-networkd-wait-online (a service whose only job is to wait until the network is up - it is what makes network-online.target from 2.14 true) at the bottom is the most common answer on a real server.
The other subcommands
systemd-analyze verify /etc/systemd/system/foo.service parse without loading
systemd-analyze security foo sandbox exposure score
systemd-analyze calendar "*-*-* 04:00:00" explain a schedule
systemd-analyze unit-paths the search path, in order
systemd-analyze dump everything (huge)
reload vs reexec vs restart
systemctl daemon-reload re-read unit files. After any edit.
systemctl daemon-reexec restart systemd's own program (PID 1) in place,
keeping all its state. After upgrading systemd.
Rare, but it is why a systemd upgrade does not
need a reboot.
systemctl restart foo restart one service. Nothing to do with the above.
Editing the whole unit
systemctl edit --full foo copies the vendor unit into /etc/systemd/system/ and opens that, so you are editing a real override rather than a drop-in. Use it when a drop-in cannot express the change - notably when you need to remove something, or reorder a list.
systemctl revert foo undoes both forms: it deletes the /etc copy and every drop-in, and the next systemctl cat shows the package's file again.
What you can now do
measure a boot and split it into kernel and userspace
find the chain of units that really delayed it, and ignore misleading blame
tell daemon-reload, daemon-reexec and restart apart
Why it helps
Boot time matters more than it seems: servers that are added automatically when load rises, or test machines created fresh for each run, are useless until they finish booting, so two extra minutes of boot cost real money and real waiting. critical-chain finds the real blocker in seconds - often a service waiting for a network interface that never comes up, or a slow disk mount.
The other subcommands are everyday tools: verify before loading any hand-written unit, security for hardening reviews, calendar before trusting a timer expression. And knowing the difference between daemon-reload, daemon-reexec and restarting a service comes up whenever someone asks "does this systemd update need a reboot?".
blame lists how long each unit took to start, slowest first, but units start in parallel. A service that took 5 seconds while nothing waited for it cost you nothing. critical-chain follows the chain of waiting up to the default target and shows which units were on the path that decided when the boot finished. Speed up the chain, not the top of the blame list.
What do @ and + mean in critical-chain?
@4.100s is how long after boot the unit became active; +87ms is how long the unit itself took to start. Units without a + took no measurable time, or are targets. Reading from the bottom up gives the order of waiting. The unit with the biggest + on the chain is your first suspect.
When do I need daemon-reexec instead of daemon-reload?
daemon-reload re-reads unit files - needed after you edit a unit. daemon-reexec restarts the systemd program itself, saving and restoring everything it knows, so a newly installed systemd version takes effect without a reboot. Package upgrades usually do it for you. You rarely type it by hand, and it is not a way to reload a service's configuration.
Why is a wait-online service so often at the end of the chain?
systemd-networkd-wait-online waits until the configured network interfaces are really up, so that network-online.target (2.14) means something. If one interface never gets a cable signal or an address, it waits until its timeout, two minutes by default, and every service that wants network-online.target waits with it. The fix is to mark interfaces that are not needed as optional in the network configuration.
What does systemd-analyze verify check?
It loads the unit files you name, plus what they refer to, without touching the running system, and reports problems: unknown or misplaced keys, bad values, relative paths, programs that do not exist, units that do not exist, and loops in the dependencies. Run it before daemon-reload. It does not run the service, so a program that crashes still needs a real start to find out.
In an interview Junior
A VM takes two minutes to boot. How do you find out why?
systemd-analyze time first: it splits the boot into the kernel part and the userspace part (the units). On a server it is almost always userspace.
Then systemd-analyze critical-chain, read bottom-up: the chain of units each waiting for the one below, with @ when each finished and + how long it took. That chain is what actually delayed the boot. The most common answer is systemd-networkd-wait-online.service, the service that makes network-online.target true, waiting for a network.
systemd-analyze blame lists units by start time, slowest first - but it misleads on its own: units start in parallel, so a slow unit nothing waits for costs nothing. Use blame to find candidates, critical-chain to find the one that matters.
Also asked: What is the difference between daemon-reload, daemon-reexec and systemctl restart? · Why can systemd-analyze blame mislead you? · What does systemctl revert undo?