systemd reports a Type=simple unit active before the service binary has been executed
reads as
`systemctl start app` returns and `systemctl is-active app` says active. Conclusion drawn: the service is up and accepting connections.
actually
For Type=simple the service manager considers the unit started immediately after the main service process has been forked off — after fork() and before the new process has called execve() to invoke the actual service binary. A unit whose binary is missing, whose port is already taken, or which needs thirty seconds to warm up, is 'active' throughout.
blind because
The process list reports existence and state. A process that will fail in a moment exists now, and readiness is simply not a quantity the manager measures for this type.
the check
Ask the socket rather than the manager: `ss -ltnp 'sport = :8000'` returns a listener only when one exists, and is empty while the unit is active but not yet serving. A single request to the port distinguishes the same two states.
cost of missing
Dependent units and deploy scripts proceed against a service that is not listening. The ordering guarantee that was assumed was never offered.
mitigation
Type=notify with sd_notify(READY=1) makes activeness mean readiness; Type=exec at least waits for execve() to succeed.
generalises to
Every start-up API that acknowledges the request rather than the readiness.
A service crash-looping every few seconds reads as active between crashes
reads as
`systemctl status app` shows active (running) with a PID. Conclusion drawn: the service is healthy.
actually
With Restart=always the unit crashes, waits RestartSec, and starts again. Sampled during a run it is active (running) with a fresh PID; sampled during the pause it is activating (auto-restart). Nothing in one sample says the PID is four seconds old and that fifty predecessors are gone.
blind because
The process list is a snapshot. A rapidly replaced process and a stable one are identical in any single frame; only the identity of the PID across frames separates them.
the check
Read the restart counter and the start timestamp twice, thirty seconds apart: `systemctl show -p NRestarts -p ExecMainStartTimestamp --value app`. A stable service returns the same two values both times; a flapping one returns different ones. Both properties are exposed by systemd for every service unit.
cost of missing
A deploy is signed off on a service that drops every request arriving inside its restart window, until the start rate limit is reached and it stays down for good.
generalises to
Every supervised process where the supervisor's diligence in restarting is read as the process's success in running.
A terminated process keeps its entry in the table until the parent reaps it
reads as
`pgrep worker` still returns a PID after the shutdown request. Conclusion drawn: the worker is refusing to exit and shutdown is hanging.
actually
The process has already exited. Its entry persists in state Z — what ps calls a 'defunct ("zombie") process, terminated but not reaped by its parent' — retaining only its exit status. It has no address space, executes nothing and cannot be killed; SIGKILL to a zombie does nothing. It disappears when the parent calls wait(), or when the parent itself exits and init reaps it.
blind because
The process list reports existence. A zombie exists as a table entry and matches by name exactly as a live process does; the state column is the only field that separates them, and name-based tools do not print it.
the check
Read the state rather than the count: `ps -o pid,stat,comm -p PID` — Z is dead, S, R or D are alive. Observed here: a forked child whose parent never waits shows STAT Z and `[python3] <defunct>`, survives `kill -9` unchanged, is matched by `pgrep python3` and is not matched by `pgrep -f`, because /proc/PID/cmdline for a zombie is 0 bytes.
cost of missing
A shutdown loop waits forever on a process that has already exited, or a supervisor counting instances by name refuses to start the replacement it should have started.
generalises to
Every registry where deregistration is a third party's responsibility: service discovery entries, connection pool slots, lock rows, task records whose owner died.
is-active reports active for a unit whose processes have all exited
reads as
`systemctl is-active provisioning` prints active. Conclusion drawn: the provisioning service is running.
actually
RemainAfterExit= 'specifies whether the service shall be considered active even when all its processes exited'. A Type=oneshot unit with RemainAfterExit=yes runs its command to completion and is then held active indefinitely with MainPID 0. Nothing is executing, and nothing will restart if the work it did is undone. Stock Ubuntu ships many such units — apparmor.service, cloud-config.service — all reading as 'active (exited)'.
blind because
is-active collapses ActiveState to a single word. That word covers both a running process and a finished one-shot, and the field that distinguishes them, SubState, is not part of the answer.
the check
`systemctl show -p SubState -p MainPID --value NAME` — `running` with a non-zero PID, or `exited` with MainPID 0. Observed here: a unit created with `systemd-run --user --property=Type=oneshot --property=RemainAfterExit=yes /bin/true` reports is-active `active`, SubState `exited`, MainPID `0` once /bin/true has returned.
cost of missing
A health check built on is-active passes for a service that finished minutes ago, and would keep passing if its binary were deleted afterwards. Restart= never fires either, because the unit is not running to fail.
generalises to
Every status vocabulary where one token spans 'in progress' and 'finished': job schedulers, CI stages, container states, queue workers reported as healthy.
A running process keeps executing a binary that has been replaced on disk
reads as
`ps -o args= -p PID` shows the daemon running from /usr/local/bin/app, and /usr/local/bin/app contains the new build. Conclusion drawn: the new build is running.
actually
Replacing an executable unlinks the old inode; a process that already mapped it holds a reference and goes on executing the previous image until it is restarted. ps prints the path recorded at exec time, which now names a different file. /proc/PID/exe still resolves, with the string '(deleted)' appended to the original pathname, and the mapped pages come from an inode with no directory entry left.
blind because
ps prints a string, not an identity. The argv and the executable path are untouched by the upgrade, so the row is byte-identical before and after it.
the check
Compare inodes rather than paths: `stat -Lc %i /proc/PID/exe` against `stat -c %i /path/to/binary`. Observed here: identical (1908455) before the upgrade; after `rm` and a fresh copy the file on disk was inode 1908456 while the process still resolved to 1908455, `readlink /proc/PID/exe` ended in '(deleted)' and /proc/PID/maps held five deleted entries. The ps output was the same in both cases.
cost of missing
A package upgrade or a deploy is verified against the file that was written while the old code keeps serving, until an unrelated restart changes the behaviour with no corresponding change to anything.
mitigation
`lsof +L1`, or `ls -l /proc/*/exe 2>/dev/null | grep deleted`, names every process still on an old image; needrestart does this after apt on Debian and Ubuntu.
generalises to
Any consumer that resolves a resource once and holds it: loaded shared libraries, opened config files, cached DNS answers, imported modules.
The process bearing the service's name is a wrapper, and the worker is its child
reads as
ps shows one process named run-analytics; it is killed and it leaves the list. Conclusion drawn: the service is stopped.
actually
A shell wrapper ending in a plain command rather than `exec` forks a child and waits on it, so there are two processes: the wrapper carrying the recognisable name, and the worker carrying the interpreter's or binary's own name. Killing the wrapper leaves the worker running, reparented to PID 1, still holding its port, its lock and its open files.
blind because
The process list is flat and reports names. Nothing in it indicates that the row matching the search is the parent of the row doing the work, and after the kill the name is genuinely gone.
the check
Verify the resource rather than the name: `ss -ltnp 'sport = :PORT'` after the stop — empty means stopped, a listener means the worker outlived the name. Observed here: killing `/bin/bash ./run-analytics` left its `sleep 120` child alive with PPID reassigned to 1; `pgrep -P PID` lists such children before the kill rather than after.
cost of missing
A restart yields two live workers competing for the same resource, or a stop that reports success while the old worker keeps writing. Neither produces an error.
mitigation
`exec` as the last line of a wrapper replaces the shell instead of forking, so the name and the worker are one process; systemd's default control-group-based KillMode stops the whole tree rather than the named process.
generalises to
Every process tree observed as a flat list: container entrypoints, npm and make targets, virtualenv shims, ssh command wrappers.
pgrep matches a fifteen-character truncation of the process name
reads as
`pgrep analytics-ingest-worker` prints nothing and exits 1. Conclusion drawn: the process is not running, so it should be started.
actually
The kernel stores a process name of at most sixteen bytes including the terminator, so ps and pgrep see 'analytics-inges'. The manual states it plainly: 'the process name used for matching is limited to the 15 characters present in the output of /proc/pid/stat'. Any pattern longer than that matches nothing, whatever is running.
blind because
The tool returns a true statement about a name it truncated, and an exit status identical to the one for 'no such process'. Newer procps prints a warning, but on stderr — the stream discarded by exactly the scripts that make this mistake.
the check
Match the command line instead: `pgrep -af analytics-ingest-worker`. Observed here with `./analytics-ingest-worker 60` running: `pgrep analytics-ingest-worker` printed nothing and exited 1, /proc/PID/comm contained 'analytics-inges', and both `pgrep -f analytics-ingest-worker` and `pgrep analytics-inges` found the process.
cost of missing
A guard that starts the service when the check fails starts a second copy of something already running: two writers on one database, two schedulers on one queue, each verified as absent immediately beforehand.
generalises to
Every lookup key silently normalised before comparison: truncated identifiers, case-folded names, unicode-normalised paths, hostnames cut at a label boundary.
A process list taken in one namespace describes a different machine
reads as
`ps aux` inside the container lists the worker as PID 1 and little else. Conclusion drawn: this is what is running on the machine, and PID 1 is the thing to signal.
actually
A PID namespace isolates a set of process IDs: a process has a different PID in each namespace it belongs to, and processes outside the namespace are invisible from within it. The listing enumerates one view. The same worker holds another PID on the host, every host process competing for the same CPU and memory is absent, and a PID copied from one view and used in the other addresses an unrelated process, if it addresses anything.
blind because
PID numbers are namespace-relative and are printed as bare integers with nothing to record which namespace produced them. Two listings of 'the processes' are simply two different sets, each internally consistent.
the check
Compare the observer's namespace with that of the process being acted on before trusting the number: `readlink /proc/self/ns/pid` against `readlink /proc/PID/ns/pid`. Identical inode strings mean the PIDs are comparable; different ones mean they are not. Observed on this host both returned `pid:[4026531836]` and `systemd-detect-virt --container` returned `none` — the case the check exists to establish rather than assume.
cost of missing
Resource accounting done inside a container attributes to itself memory the host is losing elsewhere, and a kill or restart aimed at a PID from the other view lands on whatever holds that number there.
generalises to
Every identifier unique only within a scope the output does not name: container PIDs, per-tenant row ids, per-session handles, relative paths.
free and ps report the host's memory, not the limit the process runs under
reads as
`free -h` inside the workload reports 7.8 GiB total and 4.1 GiB available. Conclusion drawn: memory is plentiful, so a slowdown or a death has some other cause.
actually
The limit is a cgroup property and /proc/meminfo is not scoped to it. memory.max is 'the main mechanism to limit memory usage of a cgroup. If a cgroup's memory usage reaches this limit and can't be reduced, the OOM killer is invoked in the cgroup.' The process may be reclaiming continuously against a ceiling two orders of magnitude below the figure free prints.
blind because
free, top and ps read /proc, which describes the machine. The constraint lives in /sys/fs/cgroup, which they do not consult.
the check
Read the limit and the pressure counters for the process's own cgroup: `CG=$(awk -F: '{print $3}' /proc/self/cgroup); cat /sys/fs/cgroup$CG/memory.max /sys/fs/cgroup$CG/memory.events`. Observed on this host inside `systemd-run --user --scope -p MemoryMax=200M`: free -h still reported 7.8Gi total and 4.1Gi available, memory.max read 209715200, and a 400 MB allocation reported success while memory.events moved from `max 0` to `max 772`, recording 772 occasions on which the limit was hit and reclaim forced.
cost of missing
Tuning is done against the wrong ceiling. A process being throttled or killed by a limit is diagnosed as slow code, and the counter that would have said so was never read.
generalises to
Every constraint enforced at a layer the inspection tool does not model: cgroup CPU quota against nproc, container disk quotas against df, API rate limits against local concurrency settings.
Load average counts processes blocked on disk as though they were running
reads as
Load average is 8.0 on a four-core machine. Conclusion drawn: the CPU is saturated, so the answer is fewer workers or more cores.
actually
proc_loadavg(5): the first three fields give 'the number of jobs in the run queue (state R) or waiting for disk I/O (state D)'. Uninterruptible sleep is counted the same as running. A load of 8 alongside an idle CPU means eight processes are blocked on storage or a stalled network filesystem, and adding cores changes nothing.
blind because
The load average is a single number covering two unlike conditions. Nothing in it separates work being done from work waiting to begin.
the check
Count the states behind the number: `ps -eo state= | sort | uniq -c`. Observed on this host at load 2.14: 101 processes in S, 74 in I, 2 in R and none in D, so the load is runnable work rather than blocked I/O. Under an I/O stall the same command shows the D column carrying the figure.
cost of missing
Capacity is added to a machine that is not short of capacity, the stall persists, and the real cause, a slow volume or a wedged mount, is never examined.
generalises to
Every aggregate that sums dissimilar states: queue depth mixing retries with new work, request counts mixing served with rejected, connection counts including half-open ones.
The %CPU column is a lifetime average, not a current rate
reads as
`ps aux` shows one process at 85.1% CPU. Conclusion drawn: this is the process consuming the machine now.
actually
ps(1) states it plainly: 'CPU usage is currently expressed as the percentage of time spent running during the entire lifetime of a process.' A process that pinned a core for six seconds and has been idle since still reports a high figure, decaying only as its lifetime grows. The converse matters more: a process running for a day that began spinning a minute ago reports a small number.
blind because
The column has the units of a rate and is read as one. Nothing in the output records the averaging window, which is each process's own age and therefore different in every row.
the check
Measure the delta over a known interval: read fields 14 and 15 of /proc/PID/stat twice and divide the difference by CLK_TCK times the elapsed seconds. Observed on this host with a process that spun for six seconds and then slept: ps reported 85.1% immediately afterwards and 22.0% twenty seconds later, while the tick delta over the following three seconds was 0 out of 300 possible, that is 0.0% actual.
cost of missing
The wrong process is restarted, throttled or blamed, and the one that has quietly started to spin is ranked below it because its long life dilutes its average.
mitigation
top's %CPU is an interval rate rather than a lifetime average, and pidstat reports per-interval figures directly.
generalises to
Every statistic whose window is implicit: lifetime averages, cumulative counters presented as gauges, uptime-normalised error rates.
kill reports success when the signal was delivered and disregarded
reads as
`kill $PID` exits zero and the deploy script moves on. Conclusion drawn: the old worker has stopped.
actually
kill(2): on success, at least one signal was sent, zero is returned. Success means the signal was queued to a process the caller had permission to signal. A process that installed an ignore disposition for SIGTERM, whether through `trap '' TERM` or a runtime that swallows it while a shutdown hook stalls, receives the signal and carries on. signal(7) notes the only exceptions: SIGKILL and SIGSTOP cannot be caught, blocked, or ignored.
blind because
The status describes the sender's half of the transaction. Whether the recipient acted is a fact about the recipient, observable only afterwards and only by looking again.
the check
Read the target's signal dispositions, or simply look again after a pause: `grep -E '^Sig(Ign|Blk|Cgt)' /proc/$PID/status`. Observed on Linux 6.8 with a script carrying `trap '' TERM` and `trap '' HUP`: two successive `kill` invocations both exited 0 and the process was still listed by `ps` after each, reporting `SigIgn: 0000000000004005`, the bits for signals 1, 3 and 15. `kill -9` ended it. `os.kill` against an unreaped zombie likewise raised nothing and returned normally.
cost of missing
The deploy continues believing the port is free. The replacement either fails to bind, or binds elsewhere and serves alongside the process that was supposed to be gone, producing a fleet where half the requests run old code.
mitigation
Treat termination as a condition to be waited on rather than an instruction to be issued: poll for the process to disappear, with a bounded escalation to SIGKILL.
generalises to
Every asynchronous request whose acknowledgement is acceptance of the message rather than performance of the work: signals, queue publishes, webhook deliveries, cache invalidations.
A unit reported inactive can still have every worker it started running
reads as
`systemctl stop app` exits zero and `systemctl is-active app` prints inactive. Conclusion drawn: the service and everything it spawned are stopped.
actually
With KillMode=process, systemd.kill(5) states that only the main process itself is killed (not recommended!), and warns that this allows processes to escape the service manager's lifecycle and resource management, and to remain running even while their service is considered stopped and is assumed to not consume any resources. The workers keep their sockets, locks and memory. The default, control-group, kills the whole cgroup and does not have this behaviour.
blind because
is-active reports the unit's state, and the unit's state is decided by its main process. Once the survivors have outlived the unit they are no longer accounted to it, so the supervisor's view is accurate and incomplete at once.
the check
Ask the kernel who is alive rather than asking systemd whether the unit is: `ps -eo pid,ppid,args | grep '[w]orker'`, or `ss -ltnp` for the port the service held. Observed on systemd 255 with a user unit `Type=simple` and `KillMode=process` whose ExecStart backgrounded a child: `systemctl --user stop` exited 0, `is-active` printed inactive, and `ps` still listed the child at pid 3917771. The identical unit at the default KillMode=control-group left nothing behind.
cost of missing
A restart appears to work while the previous generation continues serving; two versions run concurrently and diverge, and the resources the orphans hold are invisible to anything that accounts by unit.
mitigation
Leave KillMode at control-group unless there is a specific reason not to, and verify a stop by the absence of processes and listeners rather than by the unit's state.
generalises to
Any supervisor whose notion of the service is narrower than the set of processes the service created: init systems, container runtimes, CI job runners, test harnesses spawning fixtures.