Entries are organised by the instrument that missed the failure, not by the technology involved. Each states the false reading, the true state, why the instrument cannot separate them, and one discriminating check. A check qualifies only if it returns different output under the two hypotheses.
Every entry is a failure that genuinely occurs. Entries marked provenance 'observed' were diagnosed first-hand during the work that produced this site. Entries marked 'documented' cite primary documentation and their discriminating check was reproduced before publication. None are hypothetical.
The compositor-free browser reports frozen animation as no animation
reads as
A screenshot shows the element static. Conclusion drawn: the animation is badly designed, or the values are wrong.
actually
requestAnimationFrame never fires because the headless browser has no compositor. Every rAF-driven counter, canvas loop and scroll handler is frozen at frame zero, whatever its quality.
blind because
A still image cannot distinguish 'renders one frame then stops' from 'renders one frame correctly'. Both produce the same pixels.
the check
let n=0; requestAnimationFrame(()=>n++); setTimeout(()=>console.log('rAF fired:', n), 1000)
cost of missing
Every visual parameter gets tuned against a frame the animation never advances past. The tuning is not merely useless, it is fitted to an artifact.
generalises to
Any observation instrument that shares a failure mode with the thing observed.
A var assignment silently overwrites a hoisted function of the same name
reads as
A canvas is blank and unanimated. Conclusion drawn: a rendering or design problem.
actually
The file already used `var start` as a timestamp. A later `function start()` was hoisted, then overwritten by the number at execution. The call that scheduled initialisation threw a TypeError and init never ran. The element sat at its untouched default size with zero painted pixels.
blind because
An element that renders nothing and an element that renders badly both look like a design problem in a screenshot. Nothing distinguishes them visually.
the check
Read the element's backing store, not its appearance: canvas.width/height still at the 300x150 default means resize() never ran.
cost of missing
Four rounds of visual tuning applied to a layer that was never drawing.
generalises to
Any name reused across a value and a declaration in the same scope.
A relative font URL resolves one directory too deep and fails into a plausible fallback
reads as
Text renders in a serif. Conclusion drawn: the font loaded.
actually
A stylesheet at /assets/fonts.css requesting url('assets/fonts/x.woff2') resolves to /assets/assets/fonts/x.woff2 and 404s. The browser substitutes a system serif without complaint.
blind because
The fallback is a working font. Only someone who knows the intended typeface can see the substitution, and only by comparison.
the check
document.fonts.check('1em "Family Name"') or a 404 on the font path in the network log.
cost of missing
Design review proceeds against the wrong typeface. Every judgement about weight, rhythm and scale is made on a substitute.
generalises to
Every fallback that is good enough to pass inspection: default configs, cached credentials, stub implementations.
Scraping the largest asset returns a recommendation, not the subject
reads as
A parser extracts an image from the page and it is a valid, plausible image. Conclusion drawn: extraction succeeded.
actually
The page embeds related items alongside the subject. Ranking candidates by size or document order can return a neighbour, which is equally valid and equally wrong.
blind because
Both results are real images from the correct domain. Nothing about the artifact reveals it is the wrong one.
the check
Prefer the canonical marker the page declares about itself (og:image, canonical link, structured data) over any heuristic ranking of candidates.
cost of missing
Silent substitution. Detected only when two different inputs return the same output.
generalises to
Any extraction from a document that also describes things other than itself.
set -e aborts a script at a validation step that concerns something else
reads as
The install script ran and the config file is in place. Conclusion drawn: the change is active.
actually
A validation step covering the whole configuration failed on an unrelated block that needed an environment variable the script did not load. Under set -e the script exited before the reload.
blind because
The steps before the failure completed and left visible artifacts. Partial success looks like success when only the artifacts are inspected.
the check
Ask the running service what it loaded, not the filesystem what it holds. For Caddy: the admin API's live config.
cost of missing
The config is correct on disk and absent from the process, an inconsistency that survives inspection of either side alone.
generalises to
Any pipeline where a global check gates a local change.
Assets are not slow, they are queued behind synchronous work
reads as
Images take seconds to appear. Conclusion drawn: the images are too large.
actually
Their request had not been issued. Heavy synchronous work earlier in the document held the main thread, and the fetch that would load them sat unsent behind it.
blind because
Slow arrival and late departure are the same experience from the viewport.
the check
performance.getEntriesByType('resource') — startTime separates 'requested late' from 'transferred slowly'.
cost of missing
Assets get compressed, resized and lazy-loaded, which lowers quality without touching the delay.
generalises to
Any queue where wait time is read as service time.
Reveal-on-scroll renders a blank page when the observer never fires
reads as
A page loads blank in an embedded or scripted context. Conclusion drawn: a rendering failure.
actually
Elements start at opacity 0 and are revealed by an IntersectionObserver callback. Where the observer does not fire, the page is fully present and fully invisible.
blind because
The DOM is complete and correct. Only computed opacity distinguishes it from a page that failed to build.
the check
Compare element count against visible count: document.querySelectorAll('.reveal').length versus those with computed opacity above zero.
cost of missing
A working page is diagnosed as broken. Worse, the reverse: a genuinely blank page is dismissed as this.
mitigation
Any progressive-enhancement pattern that hides content by default needs a timeout that shows it regardless.
generalises to
Every design where the default state is invisible and visibility depends on a callback.
The OOM killer names the fattest process, not the one that leaked
reads as
A long-running session dies mid-task. The log records that session being killed. Conclusion drawn: that session was the problem.
actually
A different process had leaked for hours — 193 browser instances spawned by automation and never closed, 27 still resident. The OOM killer selects by current footprint, so it shot the largest process, which was an unrelated session whose context had simply grown. The leak and the casualty were different processes.
blind because
The log faithfully records the victim. It has no field for the cause, and nothing in the kill message distinguishes 'grew large' from 'made the machine run out'.
the check
Rank every process by RSS at the time of death, not just the one named: ps -eo rss,comm --sort=-rss | head -20, and count instances of anything spawned in a loop. A single fat process is a victim; a hundred medium ones are the cause.
cost of missing
The innocent session is blamed and 'fixed'. The leak keeps running and takes another process later.
mitigation
Cap or close anything spawned per-iteration, and check free memory before adding load rather than after losing work.
generalises to
Every resource-exhaustion system that reports which tenant it evicted rather than which one filled the resource.
A teardown script destroys the environment it is executing inside
reads as
Several long-running sessions vanish at once with no error output. Conclusion drawn: the tool crashed, or the machine failed.
actually
A rebuild script ran `kill-session` against the multiplexer session it was itself running in. It killed its own parent, taking four unrelated sessions with it. There is no crash and no error because the script did exactly what it was told.
blind because
A process that is killed cannot report that it was killed, and cannot report why. The absence of an error reads as an unexplained crash rather than a successful destructive command.
the check
Before any teardown, compare the target against the environment you occupy: for tmux, test whether $TMUX is set and whether its session name equals the target. Refuse if they match.
cost of missing
Work in progress across every session in the environment, lost with no diagnostic trail.
generalises to
Any tool that can destroy a container, session, service, or host that it might itself be running inside.
A privilege prompt with nowhere to appear hangs instead of failing
reads as
A deploy step produces no output and does not return. Conclusion drawn: the operation is slow, or the network is stalling.
actually
The command needed a password. There is no terminal to prompt on, so it waits indefinitely. No error, no exit code, no timeout.
blind because
Exit codes only exist for processes that exit. An instrument that reads return status has nothing at all to read, and silence resembles work in progress.
the check
Ask whether credentials are needed before running the real command: sudo -n true returns non-zero immediately when a password would be required.
cost of missing
An agent waits on a command that will never return, and a task that needed a human is reported as in progress.
mitigation
Wrap anything that might prompt in a timeout, so a hang converts into a failure you can observe.
generalises to
Every interactive prompt reached from a non-interactive context: credentials, confirmations, pagers, editors.
A pipeline returns the status of its last command, not its failing one
reads as
`npm test | tee build.log` exits zero and the log file is written. Conclusion drawn: the tests passed.
actually
A shell reports the exit status of the last command in a pipeline. The test runner exited 1; tee wrote the log and exited 0, and 0 is what the pipeline returns. `set -e` does not intervene, because the pipeline as a whole succeeded.
blind because
One number is produced for a chain of processes. The failing member's status is overwritten by its successor's, and the overwrite leaves no trace in the value the caller reads.
the check
Read the whole vector rather than the summary: `false | true; echo "${PIPESTATUS[@]}"` prints `1 0` where `$?` prints `0`. Or set `pipefail` first: `set -o pipefail; false | true` exits 1 where the same pipeline without it exits 0.
cost of missing
Every failure inside a command piped into tee, grep, jq, head or a formatter is recorded as success. A build stays green across a broken test run, and the log written alongside it is treated as proof.
mitigation
`set -euo pipefail` at the top of any script whose exit status will be believed by something else.
generalises to
Any composition that collapses several results into one and keeps the last rather than the worst.
curl exits zero after successfully downloading an error page
reads as
`curl -s -o data.json URL` exits 0 and data.json exists with content in it. Conclusion drawn: the fetch succeeded.
actually
The server answered 404 or 500. curl's task — transferring what the server chose to send — completed without fault, so the exit status is 0 and an HTML error page is now sitting in data.json under the name of the expected document.
blind because
The exit code describes the transfer, not the response. A transferred error page is a completed transfer, indistinguishable at that layer from a transferred payload.
the check
Ask for the status separately, or make curl care about it: `curl -s -o data.json -w '%{http_code}\n' URL`, or add `--fail`, which converts HTTP >= 400 into exit code 22. Observed on a 404: plain curl exits 0, `--fail` exits 22.
cost of missing
A downstream step parses an HTML error page as the config, dataset or credential file it expected. The failure surfaces at the parser, far from the request that caused it.
generalises to
Every client whose success criterion is that the protocol completed, rather than that the answer was the one asked for.
A test that does not match the discovery pattern is neither run nor reported
reads as
pytest exits 0 with a green summary after a new test is added. Conclusion drawn: the new test passes.
actually
Collection matches `test_*.py` or `*_test.py` files, and `test`-prefixed functions or methods inside `Test`-prefixed classes. A file named `tests_auth.py`, or a function named `check_expiry`, is never collected. The green result belongs entirely to the other tests. The exit code that signals an empty run, 5, applies only when nothing at all was collected, so any other test in the suite conceals the omission.
blind because
An uncollected test produces no pass line and no fail line. The summary counts what ran; it has no term for what was skipped by never being seen.
the check
`pytest --collect-only -q | grep expiry` — prints the node id if the test was collected, prints nothing if it was not. The same command distinguishes the two cases before any test is executed.
cost of missing
The behaviour the test was written to protect is unprotected, and the suite's green status is subsequently cited as evidence that it is protected.
generalises to
Any convention-driven runner where registration is implicit and non-registration is silent: test discovery, plugin loaders, autoloaded fixtures, route decorators.
A bare mock answers to method names the real object no longer has
reads as
The suite is green after a collaborator's method is renamed. Conclusion drawn: nothing depended on the old name.
actually
`Mock()` manufactures an attribute on first access and returns another Mock, which is callable and truthy. Code calling `client.charge_card(...)` against the double passes although the real class now exposes only `charge`. The test exercises an interface that no longer exists, and will keep passing however far the real object drifts.
blind because
The assertion is satisfied by the double's auto-created child. Green is a true statement about the mock, and the exit code cannot say which object the statement was about.
the check
Derive the double from the real class: `create_autospec(Client)` or `Mock(spec=Client)` raises AttributeError on exactly the call a bare `Mock()` accepted. Observed on 3.12: `Mock().exsits()` returns a truthy Mock; `create_autospec(Real).exsits()` raises AttributeError.
cost of missing
A rename is shipped with a fully green suite whose coverage of the renamed path is zero. The regression appears in production, in code the tests appeared to cover.
generalises to
Every test double whose surface is invented rather than derived from the thing it replaces.
Outside strict mode MySQL stores an adjusted value and calls the statement successful
reads as
The INSERT returns `Query OK, 1 row affected` and the client exits 0. Conclusion drawn: the row was stored as supplied.
actually
With strict mode absent from sql_mode, MySQL 'inserts adjusted values for invalid or missing values and produces warnings'. A string longer than the column is truncated to fit; `'abc'` into an integer column becomes 0. The statement is not aborted and the affected-row count is the same as for a clean insert.
blind because
Warnings are a separate channel that must be asked for. Neither the return status nor the row count changes when a value is adjusted, so the two outcomes are identical to anything reading the result of the statement.
the check
`SHOW WARNINGS` (or `SHOW COUNT(*) WARNINGS`) immediately after the statement, in the same session: it returns rows such as `Data truncated for column ...` only when a value was adjusted, and nothing when it was not.
cost of missing
Truncated identifiers and coerced numbers are indistinguishable from real data once written, and the originals are gone. Corruption is discovered by a later join that finds nothing.
mitigation
Assert the mode rather than assume it: `SELECT @@SESSION.sql_mode` should contain STRICT_TRANS_TABLES before any load is trusted.
generalises to
Any writer that repairs input rather than rejecting it: lenient parsers, schema-on-read stores, spreadsheet imports.
S3 sends 200 OK before it knows whether the upload completed
reads as
CompleteMultipartUpload returns HTTP 200. Conclusion drawn: the object is assembled and present.
actually
S3 sends the 200 header first, then keeps the connection alive with whitespace while assembly runs, which can take minutes. A failure after that point is delivered as an `<Error>` document in the body of the response whose status line already said 200. The API reference states it directly: a 200 OK response can contain either a success or an error.
blind because
The status line is written before the outcome is known, so it cannot encode the outcome. A client that reads the status and closes has read a value committed in advance of the fact it is taken to report.
the check
Parse the body even on 200 and look for an `<Error>` root element; or confirm independently with HeadObject and compare ContentLength and ETag against what was uploaded. Both differ between a completed and a failed assembly; the status code does not.
cost of missing
An upload pipeline records success for an object that does not exist. The gap is found by whatever reads it next, typically much later and in another system.
generalises to
Any protocol that must acknowledge before it can know: streamed responses, long-polling, 202-style accepted work, write-behind caches.
A batch write returns 200 while handing back the items it did not write
reads as
BatchWriteItem returns HTTP 200 and the SDK raises no exception. Conclusion drawn: all 25 items were written.
actually
The individual puts and deletes are atomic but the batch is not. Operations that failed on throughput or an internal error are returned in `UnprocessedItems` inside the 200 body, and the caller is expected to resubmit them with backoff. The low-level client hands them back; only higher-level helpers, such as boto3's batch_writer, resubmit on their own.
blind because
Total success and partial success share a status code, an exception-free return and a well-formed body. The difference is one map that is empty in the first case and populated in the second.
the check
Assert the map is empty rather than assuming it: `sum(len(v) for v in resp.get('UnprocessedItems', {}).values()) == 0`. It is 0 on a full write and non-zero whenever items were dropped.
cost of missing
Rows go missing from a bulk load in proportion to how throttled the table was, with no error recorded anywhere, and the load is reported complete.
generalises to
Every bulk endpoint that reports transport success while carrying per-item failure in its payload.
A single-page app's catch-all rewrite answers 200 for URLs that do not exist
reads as
`curl -o /dev/null -w '%{http_code}' https://site/docs/pricing` returns 200. Conclusion drawn: the page exists and the link is good.
actually
The host rewrites every unmatched path to index.html so the client-side router can handle it. The bytes returned are the application shell; the router decides only in the browser that there is nothing at this route. Google names the pattern a soft 404 and notes that such apps report 200 instead of the appropriate status code.
blind because
The status is produced by the server before any router exists. Every path under the domain, real or invented, returns the same 200 with the same shell and the same content type.
the check
Compare against a path that certainly does not exist: `curl -s $BASE/zzz-not-a-real-path | md5sum` and `curl -s $URL | md5sum`. Identical hashes mean the catch-all answered both; different hashes mean the URL has its own document.
cost of missing
Link checks, sitemap validation and 'the page is live' claims all pass against URLs with nothing behind them.
generalises to
Any fallback that answers on behalf of everything unmatched: wildcard DNS, default vhosts, permissive proxy routes.
sshd takes the first value for a keyword, so an appended directive loses to an include
reads as
/etc/ssh/sshd_config ends with `PasswordAuthentication no`, and sshd reloaded without error. Conclusion drawn: password logins are disabled.
actually
The man page states that unless noted otherwise, for each keyword the first obtained value will be used. On a stock Ubuntu image `Include /etc/ssh/sshd_config.d/*.conf` sits at line 12 of a 131-line file, so a drop-in such as 50-cloud-init.conf that sets the same keyword is read first and wins. The line appended at the bottom is parsed and discarded.
blind because
The file says what was intended, and it is the file that was edited. Precedence is a property of the merge order across several files, and the include that pre-empts the edit sits above it, out of the region being read.
the check
`sudo sshd -T | grep -i passwordauthentication` prints the effective merged value the daemon will use, which differs from the authored line whenever an earlier occurrence won. Run it with root privileges: as an unprivileged user it silently omits unreadable drop-ins.
cost of missing
A hardening change is recorded as applied while the setting it was meant to change is untouched, and the evidence for the claim is the file that lost.
generalises to
Every first-wins configuration system, which fails in exactly the opposite direction to the last-wins ones and therefore defeats the habit built on them.
Bidirectional control characters make source read differently than it compiles
reads as
A reviewer reads the diff and the early return is plainly inside a comment. Conclusion drawn: the change is inert.
actually
Unicode bidirectional overrides (U+202A to U+202E, U+2066 to U+2069) reorder the display of tokens without changing their logical order. Compilers and interpreters adhere to the logical ordering of source code, not the visual order, so the code executed is not the code rendered. Catalogued as CVE-2021-42574, with a homoglyph variant as CVE-2021-42694.
blind because
Reading a file means reading a rendering of it. The terminal, the editor and the diff viewer all apply the same bidi algorithm as the attack, so the instrument and the exploit agree with each other and disagree with the compiler.
the check
Search for the characters instead of reading the text: `grep -rlP '[\x{202A}-\x{202E}\x{2066}-\x{2069}]' path/` names files containing them and prints nothing for files that do not. Verified against a planted sample and a clean file.
cost of missing
Code is reviewed and approved on the strength of behaviour no reviewer ever saw.
mitigation
Compilers now detect this where asked: rustc's text_direction_codepoint_in_literal lint and gcc's -Wbidi-chars. Enable them rather than relying on reading.
generalises to
Any check performed on a rendering of an artifact rather than on its bytes.
A JSON integer above 2^53 is silently rounded when parsed as a double
reads as
The response contains `"id": 10765432100123456789`; the parsed object has an id of the right shape and it round-trips through the code. Conclusion drawn: the identifier was carried through intact.
actually
JavaScript parses JSON numbers as IEEE 754 doubles. The value becomes 10765432100123458000 — a different, equally plausible, non-existent identifier. RFC 8259 states that only integers within [-(2**53)+1, (2**53)-1] are interoperable in the sense that implementations will agree exactly on their values.
blind because
The corrupted value has the same type, similar magnitude and identical formatting. The sender's logs show the original and the receiver's show the rounded one, so each side is internally consistent and only a comparison across the boundary reveals the change.
the check
`Number.isSafeInteger(value)` — false for anything already rounded, true otherwise — or compare re-serialisation against the received text: `JSON.stringify(JSON.parse(s)) === s`. Verified: 10765432100123456789 parses to 10765432100123458000, isSafeInteger false, round-trip unequal; the same document parses exactly in Python.
cost of missing
Reads and writes land on the wrong record or on none. The wrongness is stable and reproducible, which makes it look like data rather than corruption.
mitigation
Carry large identifiers as strings across the boundary; APIs that learned this the hard way ship both forms, id and id_str.
generalises to
Every boundary between systems with different numeric ranges: 64-bit ids into doubles, timestamps into 32-bit seconds, decimals into floats.
systemd reports a Type=simple unit active before the service binary has been executed
reads as
`systemctl start app` returns and `systemctl is-active app` says active. Conclusion drawn: the service is up and accepting connections.
actually
For Type=simple the service manager considers the unit started immediately after the main service process has been forked off — after fork() and before the new process has called execve() to invoke the actual service binary. A unit whose binary is missing, whose port is already taken, or which needs thirty seconds to warm up, is 'active' throughout.
blind because
The process list reports existence and state. A process that will fail in a moment exists now, and readiness is simply not a quantity the manager measures for this type.
the check
Ask the socket rather than the manager: `ss -ltnp 'sport = :8000'` returns a listener only when one exists, and is empty while the unit is active but not yet serving. A single request to the port distinguishes the same two states.
cost of missing
Dependent units and deploy scripts proceed against a service that is not listening. The ordering guarantee that was assumed was never offered.
mitigation
Type=notify with sd_notify(READY=1) makes activeness mean readiness; Type=exec at least waits for execve() to succeed.
generalises to
Every start-up API that acknowledges the request rather than the readiness.
A service crash-looping every few seconds reads as active between crashes
reads as
`systemctl status app` shows active (running) with a PID. Conclusion drawn: the service is healthy.
actually
With Restart=always the unit crashes, waits RestartSec, and starts again. Sampled during a run it is active (running) with a fresh PID; sampled during the pause it is activating (auto-restart). Nothing in one sample says the PID is four seconds old and that fifty predecessors are gone.
blind because
The process list is a snapshot. A rapidly replaced process and a stable one are identical in any single frame; only the identity of the PID across frames separates them.
the check
Read the restart counter and the start timestamp twice, thirty seconds apart: `systemctl show -p NRestarts -p ExecMainStartTimestamp --value app`. A stable service returns the same two values both times; a flapping one returns different ones. Both properties are exposed by systemd for every service unit.
cost of missing
A deploy is signed off on a service that drops every request arriving inside its restart window, until the start rate limit is reached and it stays down for good.
generalises to
Every supervised process where the supervisor's diligence in restarting is read as the process's success in running.
A rotated log leaves the daemon writing to a file that no longer has a name
reads as
app.log exists, is zero bytes, and gains no lines. Conclusion drawn: the service is idle, or has stopped working.
actually
logrotate renamed or removed the file the process had open. The process still holds the old inode and keeps appending to it. The logrotate man page names the case in its description of copytruncate: it exists for programs that cannot be told to close their logfile and thus might continue writing to the previous log file forever.
blind because
Reading a log means resolving a path. After rotation the path and the process's open descriptor refer to different objects, and the reader follows the path while the writer holds the descriptor.
the check
Ask the process which file it is writing to: `ls -l /proc/$(pidof app)/fd | grep -i log`. A healthy process points at the live path; a stranded one points at a path marked `(deleted)`.
cost of missing
Log-based monitoring goes quiet and the quiet is read as calm. Disk fills with a file no directory listing can show, and it is only reclaimed when the process is restarted.
mitigation
copytruncate, or a postrotate hook that signals the daemon to reopen its log.
generalises to
Any handle held across a rename or delete: log files, config files watched by path, unlinked sockets and temp files.
Python discards records below WARNING when no logging is configured
reads as
A script instrumented with logger.info() at every step produces no output at all. Conclusion drawn: the code path never ran.
actually
With no configuration, the root logger has no handlers and the internal last-resort handler is set at WARNING. INFO and DEBUG records are created and then dropped; WARNING and above go to stderr. The code ran, and said so, into nothing.
blind because
A discarded record and a record that was never emitted produce the same empty output. A log cannot report what it filtered out, because the filtering happens before anything is written.
the check
`logging.getLogger(__name__).isEnabledFor(logging.INFO)` — False while records are being dropped, True once a handler and level are configured. Observed on 3.12: root handlers `[]`, lastResort `<_StderrHandler <stderr> (WARNING)>`, isEnabledFor(INFO) False.
cost of missing
Debugging proceeds from the false premise that the instrumented branch was not reached, and the real fault is hunted upstream of where it lives.
generalises to
Every level-filtered or sampled telemetry channel, where the absence of a line is read as the absence of an event.
A transparent overlay takes the click the screenshot shows landing on the button
reads as
The screenshot shows the button unobscured and correctly placed, and the click was dispatched without error. Conclusion drawn: the button was clicked.
actually
A transparent element — a full-viewport modal backdrop, a zero-opacity loading layer, an oversized decorative pseudo-element — covers the button's centre point, and hit testing delivers the event to the topmost element at that coordinate. WebDriver has a named error for exactly this: the Element Click command could not be completed because the element receiving the events is obscuring the element that was requested clicked.
blind because
A transparent overlay contributes no pixels. The image of a covered button and the image of an uncovered one are the same image.
the check
Ask the document what occupies the point: `const r = el.getBoundingClientRect(); document.elementFromPoint(r.left + r.width/2, r.top + r.height/2) === el` — true when the element would receive the click, false when something is over it.
cost of missing
An automated flow reports submitting forms it never submitted. A synthetic `el.click()` compounds it, because dispatching on the element directly bypasses hit testing and succeeds where a real user's click would not.
generalises to
Any interaction verified by appearance rather than by the effect the interaction was supposed to have.
A killed process loses the output it produced but never flushed
reads as
A worker is killed and its log ends several steps before the operation under investigation. Conclusion drawn: execution never reached that step.
actually
Standard output is block-buffered whenever it does not refer to a terminal, so lines accumulate in a user-space buffer of a few kilobytes until it fills or the process exits cleanly. SIGKILL cannot be caught, blocked or handled, so no flush happens. The steps ran, announced themselves, and the announcements died in the buffer.
blind because
A log records what was flushed, not what was written. A line that was never produced and a line that was produced into a buffer and then discarded are the same absence.
the check
Re-run with buffering removed and kill it the same way: `PYTHONUNBUFFERED=1 prog > out.log` (or `stdbuf -oL` for a C program). Observed on 3.12: a script printing two lines and then sleeping, SIGKILLed two seconds in, left out.log at 0 bytes; the identical run under PYTHONUNBUFFERED=1 left both lines. Output that appears only when unbuffered was being produced all along.
cost of missing
The investigation moves upstream of the last logged line, which is not where the process was. The fault lives inside the region the log appears to prove was never entered.
mitigation
Anything whose log will be read after an abnormal death should be unbuffered at the point of writing; adding it at the point of reading is too late.
generalises to
Every buffered channel inspected after an abrupt stop: stdio, log shippers with in-memory queues, metrics flushed on a timer, traces batched before export.
journald discards every message past the burst and files the notice elsewhere
reads as
`journalctl -u worker` covers the whole run and contains no error and no completion line. Conclusion drawn: the worker raised no error, and the absent completion line is the anomaly worth chasing.
actually
If more messages than RateLimitBurst are logged by a service inside RateLimitIntervalSec, all further messages within the interval are dropped until the interval is over. The default is 10000 messages in 30s, multiplied by a factor derived from the free disk space available to the journal. Everything the service says after the burst is exhausted is discarded, including the line that mattered.
blind because
The stored records are contiguous and well-formed; the log simply stops and later resumes. A message about the number of dropped messages is generated, but by journald under its own identity, so a query filtered to the unit does not show it — and a reader outside the systemd-journal and adm groups cannot see it at all.
the check
Count what the producer emitted against what the journal stored. Observed on this host: a transient unit emitting 120,001 numbered lines in 2.1s stored 37,499 of them — line 1 through line 37,499 and then nothing at all, with the final line absent and no suppression notice visible under `journalctl --user -u NAME`.
cost of missing
A verbose service is treated as a well-instrumented one, and its silence during the interesting minute is read as calm rather than as the direct consequence of its own verbosity.
mitigation
LogRateLimitIntervalSec= and LogRateLimitBurst= can be raised per unit, but a service logging at that rate needs to log less rather than louder.
generalises to
Every sampled or throttled telemetry path: metrics agents, trace sampling, syslog rate limits, ingestion quotas in hosted log services.
basicConfig does nothing once anything has already touched the root logger
reads as
`logging.basicConfig(level=logging.DEBUG)` runs at the top of main() and the program still emits only warnings. Conclusion drawn: the instrumented branches are not being reached, or the level argument is wrong.
actually
The function does nothing if the root logger already has handlers configured, unless force is set to True. One earlier call to logging.warning(), one imported library that logs during import, or one framework that configures logging first, installs a handler; every later basicConfig call is then a no-op and the root level stays at WARNING.
blind because
The call raises nothing and returns nothing to inspect, and the handler that is present is a working handler emitting real lines. The output is a correct log at the wrong level, which is far more convincing than no log at all.
the check
Interrogate the configuration rather than the output: `python3 -c 'import logging; logging.warning("x"); logging.basicConfig(level=logging.DEBUG); print(logging.root.handlers, logging.root.level, logging.getLogger().isEnabledFor(logging.INFO))'`. Observed on 3.12: handlers `[<StreamHandler <stderr> (NOTSET)>]`, level 30, isEnabledFor(INFO) False; with `force=True` the INFO record appears and isEnabledFor(INFO) is True.
cost of missing
Missing lines are attributed to unreached code, and the debugging effort goes into the application instead of into the two lines of logging setup that declined to apply.
generalises to
Every initialiser that is idempotent by doing nothing: first-wins registries, singleton bootstrappers, setup functions that check for prior state and return quietly.
A malformed log call discards its own record and returns normally
reads as
app.log holds the lines either side of a payment and no line for the payment itself. Conclusion drawn: that branch did not execute.
actually
The argument count did not match the format string. Interpolation happens inside the handler rather than at the call site, so the exception is raised during emit() and routed to Handler.handleError. The call returns normally, the program continues, the exit status is 0, and the record is gone. Where logging.raiseExceptions has been set to False — 'this is what is mostly wanted for a logging system' — nothing is written anywhere.
blind because
The line is missing for a reason internal to the logging subsystem, and the logging subsystem is the instrument being read. Its own failures are the one class of event it is built not to report through itself.
the check
Look on the other stream, which is where the default handler puts its own failures: `python3 app.py 2>&1 >/dev/null | grep -c '^--- Logging error ---'` — non-zero when records were formatted and thrown away, zero when the branch genuinely did not run. Observed on 3.12: `log.info('charged %s for %s', 'user-1')` produced a TypeError traceback on stderr, exit status 0, and an app.log containing only the following line; with `logging.raiseExceptions = False` stderr was 0 bytes and the log was identical.
cost of missing
The audit trail has a hole exactly where an operation is hardest to reconstruct, and the hole is read as evidence the operation did not happen.
mitigation
Route stderr to the same destination as the log, so the logging subsystem's own failures land beside the records they replaced.
generalises to
Every subsystem asked to report on itself: monitoring agents that cannot alert on their own death, error trackers that drop malformed events, audit logs that fail open.
Combined stdout and stderr arrive in an order that never happened
reads as
`prog > run.log 2>&1` yields three failure lines followed by three start lines. Conclusion drawn: the failures preceded the work, so something failed before the steps began.
actually
Standard output is fully buffered when it does not refer to an interactive device; standard error is not. Redirected to a file, stdout accumulates and is written in one block at exit while every stderr line goes straight through. The order in the merged file is an artefact of buffering policy, not a record of time.
blind because
A log is read as a sequence and a sequence is read as causality. The merged file records neither the originating stream nor the moment of emission, so the interleaving is the only ordering evidence available and it is precisely the part that was destroyed.
the check
Re-run with stdout unbuffered and compare the two files. Observed on 3.12: a program alternating a stdout line and a stderr line three times produced all three stderr lines before all three stdout lines under `> combined.log 2>&1`, and strictly alternating lines under `PYTHONUNBUFFERED=1`. `stdbuf -oL` did not change it, because the interpreter manages its own buffers rather than libc's.
cost of missing
A cause is assigned to the wrong step, and the fix is applied to whatever the reordering happened to place first.
mitigation
Timestamp at the point of emission, so ordering does not depend on arrival, and keep the two streams separate when their relative order carries meaning.
generalises to
Any merge of independently buffered sources into one ordered view: multi-process logs, distributed traces without synchronised clocks, tail -f across several files.
A terminated process keeps its entry in the table until the parent reaps it
reads as
`pgrep worker` still returns a PID after the shutdown request. Conclusion drawn: the worker is refusing to exit and shutdown is hanging.
actually
The process has already exited. Its entry persists in state Z — what ps calls a 'defunct ("zombie") process, terminated but not reaped by its parent' — retaining only its exit status. It has no address space, executes nothing and cannot be killed; SIGKILL to a zombie does nothing. It disappears when the parent calls wait(), or when the parent itself exits and init reaps it.
blind because
The process list reports existence. A zombie exists as a table entry and matches by name exactly as a live process does; the state column is the only field that separates them, and name-based tools do not print it.
the check
Read the state rather than the count: `ps -o pid,stat,comm -p PID` — Z is dead, S, R or D are alive. Observed here: a forked child whose parent never waits shows STAT Z and `[python3] <defunct>`, survives `kill -9` unchanged, is matched by `pgrep python3` and is not matched by `pgrep -f`, because /proc/PID/cmdline for a zombie is 0 bytes.
cost of missing
A shutdown loop waits forever on a process that has already exited, or a supervisor counting instances by name refuses to start the replacement it should have started.
generalises to
Every registry where deregistration is a third party's responsibility: service discovery entries, connection pool slots, lock rows, task records whose owner died.
is-active reports active for a unit whose processes have all exited
reads as
`systemctl is-active provisioning` prints active. Conclusion drawn: the provisioning service is running.
actually
RemainAfterExit= 'specifies whether the service shall be considered active even when all its processes exited'. A Type=oneshot unit with RemainAfterExit=yes runs its command to completion and is then held active indefinitely with MainPID 0. Nothing is executing, and nothing will restart if the work it did is undone. Stock Ubuntu ships many such units — apparmor.service, cloud-config.service — all reading as 'active (exited)'.
blind because
is-active collapses ActiveState to a single word. That word covers both a running process and a finished one-shot, and the field that distinguishes them, SubState, is not part of the answer.
the check
`systemctl show -p SubState -p MainPID --value NAME` — `running` with a non-zero PID, or `exited` with MainPID 0. Observed here: a unit created with `systemd-run --user --property=Type=oneshot --property=RemainAfterExit=yes /bin/true` reports is-active `active`, SubState `exited`, MainPID `0` once /bin/true has returned.
cost of missing
A health check built on is-active passes for a service that finished minutes ago, and would keep passing if its binary were deleted afterwards. Restart= never fires either, because the unit is not running to fail.
generalises to
Every status vocabulary where one token spans 'in progress' and 'finished': job schedulers, CI stages, container states, queue workers reported as healthy.
A running process keeps executing a binary that has been replaced on disk
reads as
`ps -o args= -p PID` shows the daemon running from /usr/local/bin/app, and /usr/local/bin/app contains the new build. Conclusion drawn: the new build is running.
actually
Replacing an executable unlinks the old inode; a process that already mapped it holds a reference and goes on executing the previous image until it is restarted. ps prints the path recorded at exec time, which now names a different file. /proc/PID/exe still resolves, with the string '(deleted)' appended to the original pathname, and the mapped pages come from an inode with no directory entry left.
blind because
ps prints a string, not an identity. The argv and the executable path are untouched by the upgrade, so the row is byte-identical before and after it.
the check
Compare inodes rather than paths: `stat -Lc %i /proc/PID/exe` against `stat -c %i /path/to/binary`. Observed here: identical (1908455) before the upgrade; after `rm` and a fresh copy the file on disk was inode 1908456 while the process still resolved to 1908455, `readlink /proc/PID/exe` ended in '(deleted)' and /proc/PID/maps held five deleted entries. The ps output was the same in both cases.
cost of missing
A package upgrade or a deploy is verified against the file that was written while the old code keeps serving, until an unrelated restart changes the behaviour with no corresponding change to anything.
mitigation
`lsof +L1`, or `ls -l /proc/*/exe 2>/dev/null | grep deleted`, names every process still on an old image; needrestart does this after apt on Debian and Ubuntu.
generalises to
Any consumer that resolves a resource once and holds it: loaded shared libraries, opened config files, cached DNS answers, imported modules.
The process bearing the service's name is a wrapper, and the worker is its child
reads as
ps shows one process named run-analytics; it is killed and it leaves the list. Conclusion drawn: the service is stopped.
actually
A shell wrapper ending in a plain command rather than `exec` forks a child and waits on it, so there are two processes: the wrapper carrying the recognisable name, and the worker carrying the interpreter's or binary's own name. Killing the wrapper leaves the worker running, reparented to PID 1, still holding its port, its lock and its open files.
blind because
The process list is flat and reports names. Nothing in it indicates that the row matching the search is the parent of the row doing the work, and after the kill the name is genuinely gone.
the check
Verify the resource rather than the name: `ss -ltnp 'sport = :PORT'` after the stop — empty means stopped, a listener means the worker outlived the name. Observed here: killing `/bin/bash ./run-analytics` left its `sleep 120` child alive with PPID reassigned to 1; `pgrep -P PID` lists such children before the kill rather than after.
cost of missing
A restart yields two live workers competing for the same resource, or a stop that reports success while the old worker keeps writing. Neither produces an error.
mitigation
`exec` as the last line of a wrapper replaces the shell instead of forking, so the name and the worker are one process; systemd's default control-group-based KillMode stops the whole tree rather than the named process.
generalises to
Every process tree observed as a flat list: container entrypoints, npm and make targets, virtualenv shims, ssh command wrappers.
pgrep matches a fifteen-character truncation of the process name
reads as
`pgrep analytics-ingest-worker` prints nothing and exits 1. Conclusion drawn: the process is not running, so it should be started.
actually
The kernel stores a process name of at most sixteen bytes including the terminator, so ps and pgrep see 'analytics-inges'. The manual states it plainly: 'the process name used for matching is limited to the 15 characters present in the output of /proc/pid/stat'. Any pattern longer than that matches nothing, whatever is running.
blind because
The tool returns a true statement about a name it truncated, and an exit status identical to the one for 'no such process'. Newer procps prints a warning, but on stderr — the stream discarded by exactly the scripts that make this mistake.
the check
Match the command line instead: `pgrep -af analytics-ingest-worker`. Observed here with `./analytics-ingest-worker 60` running: `pgrep analytics-ingest-worker` printed nothing and exited 1, /proc/PID/comm contained 'analytics-inges', and both `pgrep -f analytics-ingest-worker` and `pgrep analytics-inges` found the process.
cost of missing
A guard that starts the service when the check fails starts a second copy of something already running: two writers on one database, two schedulers on one queue, each verified as absent immediately beforehand.
generalises to
Every lookup key silently normalised before comparison: truncated identifiers, case-folded names, unicode-normalised paths, hostnames cut at a label boundary.
A process list taken in one namespace describes a different machine
reads as
`ps aux` inside the container lists the worker as PID 1 and little else. Conclusion drawn: this is what is running on the machine, and PID 1 is the thing to signal.
actually
A PID namespace isolates a set of process IDs: a process has a different PID in each namespace it belongs to, and processes outside the namespace are invisible from within it. The listing enumerates one view. The same worker holds another PID on the host, every host process competing for the same CPU and memory is absent, and a PID copied from one view and used in the other addresses an unrelated process, if it addresses anything.
blind because
PID numbers are namespace-relative and are printed as bare integers with nothing to record which namespace produced them. Two listings of 'the processes' are simply two different sets, each internally consistent.
the check
Compare the observer's namespace with that of the process being acted on before trusting the number: `readlink /proc/self/ns/pid` against `readlink /proc/PID/ns/pid`. Identical inode strings mean the PIDs are comparable; different ones mean they are not. Observed on this host both returned `pid:[4026531836]` and `systemd-detect-virt --container` returned `none` — the case the check exists to establish rather than assume.
cost of missing
Resource accounting done inside a container attributes to itself memory the host is losing elsewhere, and a kill or restart aimed at a PID from the other view lands on whatever holds that number there.
generalises to
Every identifier unique only within a scope the output does not name: container PIDs, per-tenant row ids, per-session handles, relative paths.
A truncated response arrives with its 200 already delivered
reads as
`curl -s -o data.json -w '%{http_code}' URL` prints 200 and data.json holds plausible content. Conclusion drawn: the document was retrieved.
actually
The status line is sent before the body. A connection that dies mid-body leaves a response RFC 9112 calls incomplete: 'a message body that uses the chunked transfer coding is incomplete if the zero-sized chunk that terminates the encoding has not been received', and a client that receives one 'MUST record the message as incomplete'. curl does record it, as exit code 18, but the status code stays 200 and the partial bytes stay on disk.
blind because
%{http_code} is a property of the header, which was true when it was sent. Nothing checks length, because chunked encoding exists precisely for bodies whose length is not known when the header is written.
the check
Read curl's exit status, not only the code it reports: `curl -s -o data.json -w '%{http_code}\n' URL; echo $?`. Observed against a local server that sends one chunk and closes the connection: `200` printed, exit status 18, and a 26-byte truncated prefix in data.json. `--fail` does not catch it, and `-s` suppresses the 'transfer closed with outstanding read data remaining' message that would otherwise appear.
cost of missing
A truncated CSV, NDJSON or log export is syntactically valid and merely shorter, so it loads without complaint and the missing records are indistinguishable from records that never existed.
generalises to
Every protocol that commits to a status before the payload is complete, and every consumer that validates syntax instead of completeness.
curl -I asks a different question from the one users ask
reads as
`curl -I https://site/report` returns 200 with a plausible Content-Length. Conclusion drawn: the page is up.
actually
-I sends HEAD. A server should send the same header fields it would for GET, but it is not obliged to do the same work: frameworks, proxies and CDNs commonly answer HEAD from metadata or cache without invoking the handler that renders the body. The GET a user or crawler performs can fail on the same URL at the same instant.
blind because
Status codes are per-method. A 200 to HEAD is a true statement about HEAD and says nothing whatever about GET.
the check
Request the body and discard it: `curl -s -o /dev/null -w '%{http_code}\n' URL`. Observed against a local server whose HEAD path returns metadata and whose GET path fails: HEAD 200 and GET 500 on the same URL in the same second, with `curl -I --fail` exiting 0 while `curl --fail` exited 22.
cost of missing
Uptime monitors, link checkers and post-deploy smoke tests built on -I stay green throughout an outage every real request is experiencing.
generalises to
Every cheap probe standing in for the operation being assured: a TCP connect for a request, a ping for a service, a dry run for a run.
curl drops the Authorization header when a redirect crosses to another origin
reads as
`curl -sL -u user:pass https://host/private` returns 200 with a body. Conclusion drawn: the credentials were accepted and the private resource is reachable.
actually
'Authorization:' and 'Cookie:' headers are explicitly not passed on in HTTP requests when following redirects to other origins, unless --location-trusted is used. The first hop redirected elsewhere, curl reissued the request without the credentials, and the 200 belongs to whatever that origin serves anonymously: a login page, a public shell, or an empty result set.
blind because
Every hop succeeded. The status code, the effective URL and the body are all consistent with an authenticated request, and the difference is a header removed by the client, which therefore appears in none of the server logs being consulted.
the check
Ask the final host what it received, or compare the effective URL against the origin the credentials were issued for: `curl -sL -w '%{url_effective}\n' ...`. Observed with two local servers, the second echoing the header it received: `curl -sL -u alice:secret http://127.0.0.1:8933/data` printed 'AUTH RECEIVED: None' with status 200, as did the same request carrying an explicit `-H 'Authorization: Bearer ...'`; `--location-trusted` printed the Basic credential. 127.0.0.1 and localhost count as different origins.
cost of missing
An access-control change is signed off against a response that was never authenticated. The converse is worse: the same mechanism makes a broken authentication path look like a working one.
generalises to
Every hop that rewrites a request: proxies stripping headers, SDK retries losing context, message buses dropping attributes.
A GraphQL endpoint returns 200 for a response whose data never arrived
reads as
The POST returns HTTP 200 with a JSON body of the expected shape. Conclusion drawn: the query succeeded and the body holds the data.
actually
Where the operation is executed and no request error is raised, the server should respond with 200 — 'this is the case even if a GraphQL field error is raised' during execution. Field errors are reported in an errors array while the corresponding data fields are null. Servers predating the GraphQL-over-HTTP specification answer 200 for request errors as well, so validation failures arrive the same way.
blind because
The status code describes the transport; the outcome of the operation lives in the payload. A body carrying only errors is a well-formed 200 with the correct content type.
the check
Assert on the payload: `jq -e 'has("errors") | not' resp.json`, and treat null leaves as failures rather than as absent data. Observed against the public countries.trevorblades.com endpoint: a query naming a non-existent field returned http_code 200 with a body containing only an errors array and no data entry; a valid query returned http_code 200 with a data entry.
cost of missing
A pipeline stores the null as a real value. The failure surfaces later as missing data, far from the query, with a 200 in the access log at the point where it actually went wrong.
generalises to
Every API carrying per-operation outcome in the body: JSON-RPC, batch endpoints, SOAP faults, webhook receivers that acknowledge before processing.
An unquoted YAML scalar becomes a boolean before anything reads it
reads as
config.yml reads `country: NO` and `version: 1.10`, and it is the file the service loads. Conclusion drawn: those are the values the service has.
actually
The YAML 1.1 boolean resolver matches y, yes, n, no, true, false, on and off in any capitalisation, and the loaders in widest use implement it. `NO` loads as False, `off` as False, `on` as True. Numeric resolution is equally implicit: `1.10` becomes the float 1.1 and `0xdeadbeef` becomes 3735928559.
blind because
Reading the file confirms the characters, and the characters are right. The conversion happens inside the loader, and the loaded value is never displayed beside the text it came from.
the check
Load it and print the types instead of reading it: `python3 -c "import yaml,sys;[print(repr(k),repr(v),type(v).__name__) for k,v in yaml.safe_load(open(sys.argv[1])).items()]" config.yml`. Observed on PyYAML 6.0.1: `NO` -> False, `off` -> False, `on` -> True, `1.10` -> 1.1 (float), `0xdeadbeef` -> 3735928559 (int), while `08` stayed the string '08' because it is not a valid octal literal — so neighbouring keys in one file resolve inconsistently.
cost of missing
A country code becomes a boolean and a version becomes a different version. Both are valid values of the wrong type, and both survive any review conducted by reading the file.
mitigation
Quote every scalar whose type matters. Loaders following the YAML 1.2 core schema resolve fewer of these, which changes the set of surprises rather than removing it.
generalises to
Every format that infers type from spelling: spreadsheet imports turning identifiers into dates, shell word-splitting, environment variables parsed as numbers.
A directory containing a .git is committed as a pointer rather than as files
reads as
The files are in the working tree, `git add -A` and `git commit` both succeeded, and `git status` reports a clean tree. Conclusion drawn: the files are in the repository.
actually
A directory with its own .git is recorded as a gitlink — one index entry of mode 160000 holding a commit id — not as the files beneath it. The commit id names an object in a repository nobody else can reach, and no .gitmodules entry is created. A clone contains the path as an empty directory.
blind because
git status compares the working tree with the index, and the index is satisfied: the gitlink matches the nested repository's HEAD. Everything tracked is up to date, and the untracked files sit behind a boundary status does not cross. The warning appears once, at add time, on stderr, and does not affect the exit status.
the check
Ask whether the specific file is tracked: `git ls-files --error-unmatch vendor/widget/index.js` — prints the path and exits 0 when it is, prints "did not match any file(s) known to git" and exits 1 when it is not. Observed here: `git ls-tree HEAD vendor/` returned `160000 commit bdf3631e... vendor/widget`, and a fresh clone of the repository contained README.md and nothing else.
cost of missing
The backup, the mirror or the deploy artefact is missing a subtree that every local check confirms is present. It is discovered on a clean clone, usually on another machine, usually once the original is gone.
generalises to
Every container that stores a reference where the reader assumes contents: symlinks inside archives, submodule pointers, lockfiles naming versions that no longer resolve.
The module that imports is the first file on the path with that name
reads as
config.py was edited, saved, and re-read to confirm the new value; the program still uses the old one. Conclusion drawn: a caching problem, or the process was not restarted.
actually
The interpreter searches sys.path in order, with the directory containing the input script placed at the front. Any file of that name in the script's directory, in the working directory, or earlier in the path is imported instead. The edited file is never read, and nothing is raised because the module that was found is a perfectly valid module.
blind because
The file being read and the file being imported have the same name and a similar shape, and the import statement names neither directory. Nothing in the source distinguishes them.
the check
Ask the imported module where it came from, invoked exactly as the program is invoked: `python3 -c 'import config; print(config.__file__)'`. Observed here with two config.py files present: a script in sub/ loaded sub/config.py and reported that path, while the copy that had been edited sat one directory up, untouched.
cost of missing
Edits accumulate in a file the program has never loaded, and the conclusion drawn concerns caching or process lifetime rather than identity, sending the work into restarts and clearing __pycache__.
generalises to
Every ordered resolution path where names are not unique: PATH, LD_LIBRARY_PATH, node_modules resolution, classpath, include directories.
A screenshot's pixel grid is not the page's coordinate grid
reads as
The capture shows the button with its centre at image pixel (400, 140), and a click is dispatched at (400, 140). Conclusion drawn: the click landed on the button.
actually
The capture was taken at a device scale factor above one, so image pixels and CSS pixels differ by that factor. devicePixelRatio is the ratio between the size of a device pixel and the size of a CSS pixel, and a screenshot is measured in the former while every scripting and automation coordinate is expressed in the latter. At scale 2 the button's centre is at CSS (200, 70); (400, 140) is a different part of the page.
blind because
The image carries no units. A 1600-pixel-wide PNG of an 800-pixel-wide viewport and a 1600-pixel-wide PNG of a 1600-pixel-wide viewport are both simply wide images.
the check
Compare the capture's pixel dimensions against the page's own report of its viewport: `window.innerWidth` and `window.devicePixelRatio`. Observed with Chrome 151 headless on one 800x600 window: at --force-device-scale-factor=1 the PNG was 800x600 with devicePixelRatio 1; at 2 it was 1600x1200 with devicePixelRatio 2; at 3 it was 2400x1800. The page reported an 800 CSS-pixel viewport in all three.
cost of missing
Every coordinate derived from the image is wrong by a constant factor, and the clicks land on whatever occupies the scaled position. Because something usually does, the run continues and reports the steps it believed it took.
mitigation
Take coordinates from the DOM via getBoundingClientRect, or divide image coordinates by the scale factor the capture was made at.
generalises to
Any measurement taken in one unit system and spent in another: viewport against document coordinates, physical against logical resolution, bytes against characters.
A PDF capture renders the print stylesheet rather than the page under review
reads as
The page was captured to PDF and the PDF is legible and complete. Conclusion drawn: this is what the page looks like.
actually
PDF generation switches the media type. Puppeteer states it directly: page.pdf() 'Generates a PDF of the page with the print CSS media type', and 'To generate a PDF with the screen media type, call page.emulateMediaType('screen') before calling page.pdf()'. Every @media print rule applies and every @media screen rule does not, so navigation, sticky headers and interactive affordances are commonly stripped by design.
blind because
A PDF and a screenshot are both pictures of a page, and neither records which media type produced it.
the check
Extract the text the capture actually contains and compare it against the screen render. Observed with Chrome 151 headless on a page carrying visible text in a .screen-only element plus `@media print{.screen-only{display:none} body::after{content:'PRINT STYLES ACTIVE'}}`: --print-to-pdf produced a file whose only text-showing operators decoded to 'PRINT STYLES ACTIVE'. The words 'Screen layout' appear nowhere in it.
cost of missing
A layout is signed off against an artifact no visitor will ever see, and print rules written long beforehand, often to remove exactly the elements being reviewed, silently define the record.
generalises to
Any render whose conditions are chosen by the renderer rather than the document: print media, forced colours, reduced motion, emulated devices.
A headless capture exercises one branch of a colour-scheme fork
reads as
The screenshot shows the page correctly styled and legible throughout. Conclusion drawn: the page renders correctly.
actually
prefers-color-scheme resolves to a single value per render, and a headless browser with no desktop session reports light. Every rule inside `@media (prefers-color-scheme: dark)` was parsed, matched nothing and contributed no pixels. The dark render, which a large share of visitors receive, was never produced at all.
blind because
A screenshot is one render under one set of resolved media features. The branch that did not match leaves no trace in the image, so 'the dark theme is correct' and 'the dark theme was never evaluated' look identical.
the check
Ask the page which branch it is in, and capture both: `matchMedia('(prefers-color-scheme: dark)').matches`. Observed with Chrome 151 headless: the default run reported dark=false, light=true; the same page under --force-dark-mode reported dark=true, light=false. The page also reported prefers-reduced-motion and forced-colors as inactive by default, so those branches are unrendered for the same reason.
cost of missing
Contrast failures, unreadable text and unstyled surfaces ship in the branch nobody rendered, and remain invisible to every subsequent screenshot taken the same way.
generalises to
Every conditional whose condition is supplied by the environment: feature flags defaulting off, locale-dependent formatting, reduced-motion and forced-colours branches.
A capture taken at the load event shows the designed empty state
reads as
The screenshot shows a clean, well-styled page reading 'No results'. Conclusion drawn: the query returned nothing, so the filter or the data is wrong.
actually
The load event fires once the document and its declared subresources have loaded. It says nothing about fetches started by scripts. The request was still in flight, so the placeholder provided for a genuinely empty result was the thing on screen.
blind because
The empty state is a real, intentional, correctly styled view. A picture of 'no data yet' and a picture of 'no data at all' are the same picture, because the same markup produced both.
the check
Count the data-bearing elements at capture time instead of judging the image: `document.querySelectorAll('#list li').length`. Observed against a local endpoint delayed by two seconds: the load-event capture reported rows=0 with the empty state visible, while a capture taken after the fetch resolved reported data-rows=2 and contained `<li>alpha` and `<li>beta`. The two PNGs differed in 1,262 pixels.
cost of missing
Investigation moves to the query, the filter and the backend, none of which are broken. The moment the picture was taken is the one variable never questioned.
mitigation
Trigger the capture on an assertion about content rather than on load; a network-idle condition is weaker but still better than the load event.
generalises to
Every designed representation of absence: empty tables, zero counts, blank dashboards, 'no alerts' panels.
wait without arguments returns zero however its children exited
reads as
A script starts several jobs with `&`, calls `wait`, and exits zero. Conclusion drawn: every job succeeded.
actually
The bash manual is explicit: 'If id is not given, wait waits for all running background jobs and the last-executed process substitution, if its process id is the same as $!, and the return status is zero.' The children's statuses are reaped and discarded. Any number of them may have failed.
blind because
An exit code reports what the last command chose to return, and bare wait returns zero by specification. The failing work happened in processes whose status was never requested.
the check
Wait on each recorded PID and keep the statuses: `rc=0; for p in "${pids[@]}"; do wait "$p" || rc=$?; done; exit $rc`. Observed on bash 5.2.21 with one child exiting 3 and another exiting 7: bare `wait` returned 0, `wait $pid` on the second returned 7, and `wait -n` returned the status of the first job to finish.
cost of missing
Parallelism converts a failing step into a silent one. A fan-out of uploads, migrations or builds reports success while an arbitrary subset of it did not happen.
generalises to
Every aggregator that reduces many statuses to one: parallel test runners, job schedulers, batch APIs, fan-out without per-item bookkeeping.
Output redirection empties the file before the command reads it
reads as
`sort data.txt > data.txt` exits zero and data.txt is still there. Conclusion drawn: the file was sorted in place.
actually
The shell performs redirections before running the command, and for output redirection 'if the file does not exist it is created; if it does exist it is truncated to zero size'. sort then opens an empty file, reads nothing and writes nothing, correctly and successfully. The original contents are gone.
blind because
The exit code belongs to a command that did exactly what was asked of it with the input it was given. The destruction happened in the shell, before the command started, and produced no status of its own.
the check
Compare the line count before and after in the same command. Observed on bash 5.2.21: a three-line data.txt held zero lines after `sort data.txt > data.txt`, with sort exiting 0; `grep -v DEBUG conf.txt > conf.txt` left conf.txt at zero bytes, with grep exiting 1 because it had nothing to match.
cost of missing
The file that was supposed to be filtered is now empty, and emptiness is valid input to whatever reads it next: configuration becomes all-defaults, a dataset becomes zero records, and neither state raises an error.
mitigation
Write to a new name and rename over the original, or use a tool with an explicit in-place mode such as `sed -i` or `sponge`.
generalises to
Every operation that opens its destination before reading its source: in-place archive rewrites, dumps piped over their own file, copies where source and destination alias.
find's exit status describes find's own traversal. The manual says it 'exits with status 0 if all files are processed successfully, greater than 0 if errors occur' and calls this 'deliberately a very broad description'. The status of each -exec child is not part of it. All of them may have failed.
blind because
One process walked the tree and a different process did the work. The status available to the caller belongs to the one that only walked.
the check
Dispatch through a tool whose status covers the children: `find . -type f -print0 | xargs -0 -n1 validate`, which exits 123 'if any invocation of the command exited with status 1-125'. Observed on GNU findutils with two matching files: `find f -type f -exec false \;` exited 0 and `-exec sh -c 'exit 3' \;` also exited 0, while `find f -type f -print0 | xargs -0 -n1 false` exited 123 and the same pipeline with `true` exited 0.
cost of missing
A validation, conversion or upload sweep reports success across an entire tree while every item in it failed, and a sweep is precisely the step nobody re-checks item by item.
generalises to
Every dispatcher whose status covers dispatch rather than outcome: cron wrappers, CI steps that shell out, message producers acknowledged on enqueue.
A source path without a trailing slash adds a directory level at the destination
reads as
`rsync -a build /var/www/site/` exits zero and the files are present under /var/www/site. Conclusion drawn: the build was deployed.
actually
rsync's manual: 'A trailing slash on the source changes this behavior to avoid creating an additional directory level at the destination.' Without it the directory is copied by name, so the files land in /var/www/site/build/, one level below where the server is configured to look. The previous build continues to be served.
blind because
The transfer succeeded and every file was copied correctly to a real path. The exit code describes the copy, not the destination's relationship to whatever reads it.
the check
List the destination rather than trusting the status: `find /var/www/site -maxdepth 2 -name index.html`. Observed on rsync 3.2.7: `rsync -a rs/src rs/dest/` exited 0 and produced rs/dest/src/index.html, while `rsync -a rs/src/ rs/dest/` exited 0 and produced rs/dest/index.html.
cost of missing
The deploy reports success and the site does not change. Re-running it reproduces the same success, so the natural response to the symptom confirms the wrong hypothesis.
generalises to
Every copy whose destination semantics depend on a trailing character or on whether the target already exists: cp, scp, docker COPY, object-storage prefixes.
A response saved without decompression is stored as its compressed bytes
reads as
`curl -H 'Accept-Encoding: gzip' -o data.json URL` reports 200, exits zero, and data.json is the expected size. Conclusion drawn: the document was fetched.
actually
curl decompresses only when it negotiated the encoding itself. --compressed 'Request[s] a compressed response using one of the algorithms curl supports, and automatically decompress[es] the content'; a hand-written Accept-Encoding header asks for gzip without arranging for it to be undone. The file on disk begins 1f 8b and is a gzip member, not JSON.
blind because
Status, exit code and transferred byte count are identical to a successful plain fetch. %{size_download} counts wire bytes, so it agrees under both hypotheses: 67 bytes either way.
the check
Ask what the file is rather than how big it is: `file -b data.json`. Observed on curl 8.5.0 against a local gzip-encoding server: with a hand-set header the file was 'gzip compressed data' and `grep -c alpha data.json` found no match and exited 1; with --compressed the same URL produced 'JSON text data' and the same grep printed 1.
cost of missing
A search over the artifact returns nothing and the absence is read as a fact about the content. Tools that treat the body as opaque, such as archivers, uploaders and checksums, propagate the encoded bytes without ever failing.
mitigation
curl's manual warns that saved response headers are not modified, so a stored header still claims the content is compressed after curl has decompressed it. The header is not a reliable record either way.
generalises to
Every transport-level transformation the receiver must undo: content-encoding, transfer-encoding, base64 envelopes, client-side decryption.
A 200 carrying an Age header was answered by a cache, not by the origin
reads as
`curl -sI https://site/` returns 200 and a Date header. Conclusion drawn: the origin is serving this, now.
actually
RFC 9111: 'When a stored response is used to satisfy a request without validation, a cache MUST generate an Age header field, replacing any present in the response with a value equal to the stored response's current_age.' A nonzero Age is a shared cache stating how old the body is, and the Date header records when that stored response was generated rather than when the request was made.
blind because
A status code does not identify the responder. Every hop returns 200 on a cache hit, and the body is a complete, valid, previously correct document.
the check
Read the caching headers alongside the status: `curl -sI URL | grep -iE '^(date|age|x-cache|cf-cache-status):'`. Observed at 20:07:28 UTC: https://vercel.com/ returned `age: 551` with `date: Sun, 23 Aug 2026 19:58:15 GMT`, a body nine minutes old, under `cache-control: public, max-age=0, must-revalidate`; https://developer.mozilla.org/ returned `age: 3066` with `x-cache: MISS, HIT, HIT`; a Cloudflare-fronted origin returned `cf-cache-status: DYNAMIC` and no Age at all.
cost of missing
A deploy is verified against content that predates it, and the verification is stable: repeating the request returns the same stored copy with a larger Age, which reads as consistency.
mitigation
A cache-busting query string proves the origin is correct but says nothing about what visitors receive. Both readings are needed, and they answer different questions.
generalises to
Every layer that may answer on another's behalf: CDNs, reverse proxies, resolvers, package mirrors, SDK-level response caches.
A request issued from the origin host does not travel the path visitors take
reads as
`curl -s -o /dev/null -w '%{http_code}' https://example.com/` run on the server returns 200. Conclusion drawn: the site is reachable and correct for visitors.
actually
The name resolved to an address that short-circuits the public path: an /etc/hosts entry, a split-horizon resolver, or the machine's own public address. The origin answered directly, and the CDN, WAF, redirect rules and edge certificate that every visitor traverses were not involved. A failure in any of them is unreachable by this request.
blind because
A status code records that something answered. Which of several layers answered is encoded nowhere in it, and a healthy origin behind a broken edge returns the same 200 as a healthy edge.
the check
Record who answered, not just what: `curl -s -o /dev/null -w 'code=%{http_code} remote=%{remote_ip}\n' URL`. Compare that address against the origin you deployed to. A proxied domain answers from the proxy's address whether or not the origin behind it is alive; an unproxied one answers from the origin itself. Only the second reading tells you the origin is serving.
cost of missing
Edge misconfiguration, an expired certificate, a wrong origin rule or a route blocked by a WAF is confirmed working by a check structurally incapable of reaching it.
generalises to
Any probe issued from inside the system it measures: internal health checks, same-network monitoring, tests that resolve names through a private zone.
git add says nothing when a pathspec matches only ignored files
reads as
`git add .` exits zero, the commit succeeds, and `git status` afterwards reports a clean tree. Conclusion drawn: everything in the working directory is committed.
actually
git-add documents both branches: 'The git add command will not add ignored files by default. You can use the --force option to add ignored files. If you specify the exact filename of an ignored file, git add will fail with a list of ignored files. Otherwise it will silently ignore the file.' A broad pathspec takes the silent branch, and `git status` does not list ignored files, so the tree reads as clean.
blind because
Both instruments agree, and both are answering a narrower question than the one asked. `git status` compares the index against the tracked working tree; a file that is neither tracked nor reportable is outside that comparison by construction.
the check
Ask whether a specific path is excluded, and list what was excluded: `git check-ignore -v path` and `git status --short --ignored`. Observed on git 2.43.0 with a .gitignore containing dist/, *.local and config/*: `git add .` exited 0, `git status --short` listed only .gitignore and app.py, `git ls-files` confirmed two tracked files, and `git status --short --ignored` printed `!! config/`, `!! dist/` and `!! settings.local` for the three that were never staged.
cost of missing
A build output, a migration or a generated asset that a broad rule happens to match is absent from every clone and every deploy, and each step in the chain reports the tree as clean.
generalises to
Every filter applied before a report is produced: exclude rules in backups and syncs, packaging manifests, .dockerignore, test collection patterns.
Two readers of one CRLF file disagree about where each value ends
reads as
cat .env prints `API_URL=https://api.example.com` and that is the file the service loads. Conclusion drawn: the service has that URL.
actually
The file uses CRLF terminators. Python's default text mode translates them, so a Python reader sees a 23-character value; the shell does not translate, so `. ./.env` yields a 24-character value ending in a carriage return. The same bytes become different strings depending on who reads them.
blind because
Reading the file with cat, an editor or a code review renders the carriage return as nothing at all. It has no glyph, occupies no column, and is removed by several of the tools used to inspect it.
the check
Make the terminators visible, or measure the value in the reader that matters: `cat -A .env`. Observed on this host: cat -A printed `API_URL=https://api.example.com^M$`, `file` reported 'ASCII text, with CRLF line terminators', bash reported ${#API_URL} as 24 and the equality test against the intended URL failed, while Python text mode reported 23 and the same file opened in binary mode yielded 'https://api.example.com\r'.
cost of missing
A hostname, token or path acquires an invisible trailing byte. Comparisons fail, signatures do not verify, and a request goes to a name that differs from the one on screen, while every review of the file keeps confirming the value is right.
mitigation
Normalise on ingest with a gitattributes `text` rule, and compare lengths rather than appearances when a value refuses to match something it visibly equals.
generalises to
Every character that renders as nothing: byte-order marks, zero-width spaces, non-breaking spaces pasted from documents, trailing whitespace in secrets.
Copying a symlink in archive mode produces a second link, not a backup
reads as
`cp -a app.conf app.conf.bak` exits zero and a listing shows both names. Conclusion drawn: the original is preserved, so the edit is safe.
actually
Archive mode implies --no-dereference and --preserve=links: symbolic links are copied as symbolic links. app.conf was a link, so app.conf.bak is a second link to the same target. There is one file. Editing through either name changes both, and the backup records nothing.
blind because
Reading either path returns the intended contents, and a listing shows two entries with the expected names. Only the link marker and the inode number distinguish a backup from an alias.
the check
Compare inodes after dereferencing: `stat -Lc '%i %n' app.conf app.conf.bak`. Observed on GNU coreutils: after `cp -a app.conf app.conf.bak` both names and the underlying real.conf reported inode 2142629, and overwriting app.conf with new content changed the contents visible through app.conf.bak at the same moment.
cost of missing
The rollback path does not exist, and its absence is discovered only when it is needed. A policy requiring a backup before editing is satisfied on paper by an operation that made none.
mitigation
`cp -L` copies the target's contents. Checking the inode immediately afterwards costs one command and is the only cheap moment to find this.
generalises to
Every duplication that may preserve a reference instead of the data: hard links, copy-on-write clones, container image layers, object-store copies that alias.
free and ps report the host's memory, not the limit the process runs under
reads as
`free -h` inside the workload reports 7.8 GiB total and 4.1 GiB available. Conclusion drawn: memory is plentiful, so a slowdown or a death has some other cause.
actually
The limit is a cgroup property and /proc/meminfo is not scoped to it. memory.max is 'the main mechanism to limit memory usage of a cgroup. If a cgroup's memory usage reaches this limit and can't be reduced, the OOM killer is invoked in the cgroup.' The process may be reclaiming continuously against a ceiling two orders of magnitude below the figure free prints.
blind because
free, top and ps read /proc, which describes the machine. The constraint lives in /sys/fs/cgroup, which they do not consult.
the check
Read the limit and the pressure counters for the process's own cgroup: `CG=$(awk -F: '{print $3}' /proc/self/cgroup); cat /sys/fs/cgroup$CG/memory.max /sys/fs/cgroup$CG/memory.events`. Observed on this host inside `systemd-run --user --scope -p MemoryMax=200M`: free -h still reported 7.8Gi total and 4.1Gi available, memory.max read 209715200, and a 400 MB allocation reported success while memory.events moved from `max 0` to `max 772`, recording 772 occasions on which the limit was hit and reclaim forced.
cost of missing
Tuning is done against the wrong ceiling. A process being throttled or killed by a limit is diagnosed as slow code, and the counter that would have said so was never read.
generalises to
Every constraint enforced at a layer the inspection tool does not model: cgroup CPU quota against nproc, container disk quotas against df, API rate limits against local concurrency settings.
Load average counts processes blocked on disk as though they were running
reads as
Load average is 8.0 on a four-core machine. Conclusion drawn: the CPU is saturated, so the answer is fewer workers or more cores.
actually
proc_loadavg(5): the first three fields give 'the number of jobs in the run queue (state R) or waiting for disk I/O (state D)'. Uninterruptible sleep is counted the same as running. A load of 8 alongside an idle CPU means eight processes are blocked on storage or a stalled network filesystem, and adding cores changes nothing.
blind because
The load average is a single number covering two unlike conditions. Nothing in it separates work being done from work waiting to begin.
the check
Count the states behind the number: `ps -eo state= | sort | uniq -c`. Observed on this host at load 2.14: 101 processes in S, 74 in I, 2 in R and none in D, so the load is runnable work rather than blocked I/O. Under an I/O stall the same command shows the D column carrying the figure.
cost of missing
Capacity is added to a machine that is not short of capacity, the stall persists, and the real cause, a slow volume or a wedged mount, is never examined.
generalises to
Every aggregate that sums dissimilar states: queue depth mixing retries with new work, request counts mixing served with rejected, connection counts including half-open ones.
The %CPU column is a lifetime average, not a current rate
reads as
`ps aux` shows one process at 85.1% CPU. Conclusion drawn: this is the process consuming the machine now.
actually
ps(1) states it plainly: 'CPU usage is currently expressed as the percentage of time spent running during the entire lifetime of a process.' A process that pinned a core for six seconds and has been idle since still reports a high figure, decaying only as its lifetime grows. The converse matters more: a process running for a day that began spinning a minute ago reports a small number.
blind because
The column has the units of a rate and is read as one. Nothing in the output records the averaging window, which is each process's own age and therefore different in every row.
the check
Measure the delta over a known interval: read fields 14 and 15 of /proc/PID/stat twice and divide the difference by CLK_TCK times the elapsed seconds. Observed on this host with a process that spun for six seconds and then slept: ps reported 85.1% immediately afterwards and 22.0% twenty seconds later, while the tick delta over the following three seconds was 0 out of 300 possible, that is 0.0% actual.
cost of missing
The wrong process is restarted, throttled or blamed, and the one that has quietly started to spin is ranked below it because its long life dilutes its average.
mitigation
top's %CPU is an interval rate rather than a lifetime average, and pidstat reports per-interval figures directly.
generalises to
Every statistic whose window is implicit: lifetime averages, cumulative counters presented as gauges, uptime-normalised error rates.
Redirection order decides whether the log can contain errors at all
reads as
The service runs as `app 2>&1 > app.log` and app.log holds a clean sequence of startup lines with no errors. Conclusion drawn: the run was clean.
actually
Redirections are processed left to right. The bash manual gives this exact pair: `ls > dirlist 2>&1` sends both streams to the file, while `ls 2>&1 > dirlist` 'directs only the standard output to file dirlist, because the standard error was duplicated from the standard output before the standard output was redirected to dirlist'. Standard error went wherever standard output pointed beforehand, usually a terminal that no longer exists or a parent's discarded output.
blind because
The log is genuine, complete and correctly ordered for the stream it captured. Nothing in it can indicate that a second stream existed and went elsewhere, so a log with no errors and a log that cannot contain errors are the same file.
the check
Ask the running process where its descriptors point: `readlink /proc/$$/fd/1 /proc/$$/fd/2` from inside the redirected command. Observed on bash 5.2.21: under `./probe.sh 2>&1 > out1.log`, fd 1 pointed at out1.log while fd 2 pointed at the parent's output; under `./probe.sh > out2.log 2>&1` both pointed at out2.log. A script emitting one error line produced `grep -c ERROR` of 0 in the first case and 1 in the second.
cost of missing
Every diagnostic the program emits is discarded by the same arrangement that produces the record used to declare it healthy, and alerting built on that log's contents can never fire.
generalises to
Any capture configured to watch one channel while the interesting events use another: stderr against stdout, structured logs against panics, application logs against the supervisor's.
A service's stdout is recorded at info priority whatever the line says
reads as
`journalctl -u app -p err` prints '-- No entries --'. Conclusion drawn: the service has logged no errors.
actually
systemd assigns the priority, not the text. SyslogLevel= is 'the default syslog log level to use when logging to the logging system or the kernel log buffer', it 'only applies to log messages written to stdout or stderr', and it 'Defaults to info'. Unless a line carries an explicit angle-bracket level prefix, every line the process prints is stored at priority 6, including the ones whose text reads ERROR.
blind because
The filter and the store agree. Priority is metadata attached at ingestion, so a severity filter reports the transport's opinion rather than the application's.
the check
Look at the priority distribution instead of the filtered view: `journalctl -u UNIT -o json | jq -r .PRIORITY | sort | uniq -c`. Observed on a unit logging 43,061 records over three weeks: every one at PRIORITY 6 (informational), while a plain-text search of the same range found 14 lines containing 'error'. `journalctl -u UNIT -p err` reported '-- No entries --' throughout.
cost of missing
Severity-based alerting and triage are silently disabled for every service that logs to stdout without prefixes, which is most of them, and the absence of high-priority records is read as the absence of high-priority events.
mitigation
Emit the angle-bracket level prefix from the application, or set SyslogLevel= on the unit; SyslogLevelPrefix= controls whether such prefixes are honoured.
generalises to
Every field assigned by a collector rather than by the source: levels inferred by a log shipper, statuses rewritten by a proxy, timestamps stamped at ingestion.
A journal with no persistent directory discards its evidence at reboot
reads as
After a crash and a restart, `journalctl -u app --since '2 days ago'` returns nothing. Conclusion drawn: the service logged nothing before it died, so the failure was abrupt.
actually
journald's Storage= defaults to auto, and 'auto behaves like persistent if the /var/log/journal directory exists, and volatile otherwise (the existence of the directory controls the storage mode)'. Under volatile storage the journal lives below /run and does not survive a reboot. The pre-crash records existed, and the restart performed to recover deleted them.
blind because
An empty query result has one shape. 'Nothing was logged', 'nothing matched the filter' and 'the storage that held it no longer exists' all render as no output and a zero exit.
the check
Establish whether history survives before drawing conclusions from its absence: `ls -d /var/log/journal 2>/dev/null; journalctl --list-boots`. Observed on this host: /var/log/journal exists and holds 566 MB, and --list-boots lists two boots reaching back five weeks, so an empty result here is a fact about the service. On a host without that directory the same commands print nothing and a single boot, and no empty result carries information.
cost of missing
The post-mortem proceeds from the premise that the process died silently, and the reboot performed to restore service is the act that destroyed the evidence for any other explanation.
mitigation
Create /var/log/journal and restart systemd-journald, or set Storage=persistent explicitly, before the next incident rather than after it.
generalises to
Every store whose retention is shorter than the investigation: ring buffers, in-memory metrics, container logs removed with the container, tmpfs working directories.
Traffic classified by user-agent counts what clients claim to be
reads as
An access log shows 57% of requests from browsers. Conclusion drawn: most visitors are people, and the site is reaching a human audience.
actually
The User-Agent header is set by the client and asserted, never verified. A single scanner sending a stock Windows Chrome string produced a large share of that bucket; its requests were for /contact, /about-us, /pricing, /team and /support — pages this site has never had. Real browsers and anything imitating one are indistinguishable by header alone.
blind because
The field being counted is supplied by the party being measured. Every row is internally consistent and none of them is evidence.
the check
Group requests by client address and compare what each one asked for against what exists. A client whose requests are mostly 404s for pages the site has never published is enumerating, whatever it calls itself. Corroborate with an independent signal the client does not control, such as whether it also fetched the page's own subresources.
cost of missing
Audience is misread in the direction that flatters. Content decisions get made for readers who were never there, and genuine machine traffic is filed as human.
mitigation
Treat the user-agent as one weak signal among several. Behaviour — which paths, in what order, with which subresources — is set by the client too, but it is far more expensive to fake convincingly.
generalises to
Any metric derived from a field the measured party supplies: referrers, self-reported versions, declared content types, client-side analytics events.
A capture is sized to the document, so horizontal overflow has nowhere to show
reads as
The full-page capture shows every section filling the frame, with nothing clipped at either edge and no scrollbar anywhere in the image. Conclusion drawn: the layout fits the viewport.
actually
Viewport-percentage units ignore scrollbars. CSS Values 4 is explicit: the viewport-percentage lengths are sized assuming that scrollbars do not exist, even if this diverges from the initial containing block. On a window 800 CSS pixels wide with a classic 15-pixel scrollbar, percentage widths resolve against 785 while 100vw resolves to 800, so every full-bleed element overhangs the layout by exactly one scrollbar and the document acquires a horizontal scrollbar of its own.
blind because
A screenshot has no scrollbars and no edges beyond the content. The capture is made as wide as the document's scroll width, so the overflowing strip is inside the image rather than past its edge, and the image is the same image it would be if nothing overflowed.
the check
Ask the document whether it is wider than its own viewport: `document.documentElement.scrollWidth - document.documentElement.clientWidth`. Observed with Chrome 151 headless at --window-size=800,600 on a vertically overflowing page: innerWidth 800, clientWidth 785, a `width:100vw` box measured 800px, a `width:100%` box measured 785px, and the difference came back as 15. On the same browser a page without any 100vw element produced a 785-pixel-wide full capture; the page with one produced an 800-pixel-wide capture, neither showing a clipped edge.
cost of missing
Every desktop visitor gets a horizontal scrollbar on every page, and on touch devices the page rubber-bands sideways. The defect is reported by users and cannot be reproduced from any capture, which sends the investigation to the wrong layer.
mitigation
Size full-bleed elements against the containing block rather than the viewport, or reserve the gutter with `scrollbar-gutter: stable` so the two frames of reference agree.
generalises to
Any measurement taken in a frame that excludes the thing being measured: viewport units against layout width, container queries against a resized container, timings taken inside the operation being timed.
A full-page capture ends where the renderer decided to stop rendering
reads as
The full-page capture is 2406 pixels tall, every section in it is drawn, and the article reads through to its end. Conclusion drawn: this is the whole page.
actually
The sections carry `content-visibility: auto`, which CSS Containment 2 defines as turning on layout, style and paint containment and, if the element is not relevant to the user, also skipping its contents. Skipped contents are not painted, as if they had visibility: hidden, and the element is sized by its `contain-intrinsic-size` placeholder instead of by what is inside it. Offscreen sections therefore contribute their placeholder height to the document and nothing to the image.
blind because
A capture records what was painted. A subtree that was never laid out contributes no pixels and no height, so a short document and a truncated one are the same picture: complete, continuous and ending in a plausible place.
the check
Force the skipping off and re-measure the document: `(() => { const a = document.documentElement.scrollHeight; document.querySelectorAll('*').forEach(e => e.style.contentVisibility = 'visible'); return [a, document.documentElement.scrollHeight]; })()`. Observed with Chrome 151 headless on a page with three `content-visibility: auto` sections declaring `contain-intrinsic-size: auto 300px` around 718 pixels of real content each: [2406, 3654]. Each section measured 302px rather than 718px, and 1248 pixels of article were absent from the layout and from the capture alike.
cost of missing
Content review, screenshot diffing and visual regression all operate on a document that is missing most of itself, and the missing part is the part nobody scrolled to, which is where unreviewed content accumulates.
mitigation
Set `contain-intrinsic-size` to a value close to the real height so the placeholder does not distort the document, and disable content-visibility for any automated capture.
generalises to
Every optimisation that does less work when nobody is watching: lazy loading, virtualised lists, deferred hydration, sampled tracing.
A frame the embedded site refused renders as ordinary whitespace
reads as
The capture shows the dashboard with a clean empty band where the third-party widget sits, and the DOM confirms the iframe is present with the right src. Conclusion drawn: the widget loaded and has nothing to display.
actually
The embedded document declined to be framed. RFC 7034 on X-Frame-Options: DENY means a browser receiving content with this header field MUST NOT display this content in any frame. The iframe element is still laid out at its declared size and left empty. Nothing about the parent document changes, and no layout shift marks the refusal.
blind because
An iframe reserves its box before it has any content, so a frame that was refused and a frame that rendered a blank empty state occupy the same rectangle of the same colour. Inspecting the DOM confirms the element and its src, both of which are correct; what failed is on the other side of the boundary.
the check
Ask the resource timeline what arrived rather than the DOM what exists: `performance.getEntriesByType('resource').filter(e => e.initiatorType === 'iframe').map(e => e.name + ' ' + e.transferSize)`. Observed with Chrome 151 headless against a local server: the refused frame reported transferSize 0 and logged `Refused to display 'http://127.0.0.1:8936/' in a frame because it set 'X-Frame-Options' to 'deny'`; the identical page pointed at an unprotected copy reported transferSize 405. `frames.length` was 1 and the iframe's src was the intended URL in both runs, and counting pixels inside the frame's rectangle gave 0 widget-coloured pixels against 107,776.
cost of missing
A payment form, a status board or a support widget is absent for every visitor while every capture and every DOM assertion says it is there. Because the parent page is intact, monitoring built on the parent stays green.
mitigation
Treat an embed as a dependency with its own health check: assert on the frame's load event or its resource entry, not on the presence of the element.
generalises to
Every boundary where the failure is declared by the far side and absorbed silently by the near one: blocked embeds, CORS-rejected fetches, refused redirects, sandboxed scripts.
An ancestor's overflow reassigns what a sticky element sticks to
reads as
The header is declared `position: sticky; top: 0`, the capture shows it at the top of the page, and getComputedStyle reports `sticky`. Conclusion drawn: it sticks.
actually
CSS Position 3 defines sticky as identical to relative except that its offsets are automatically adjusted in reference to the nearest ancestor scroll container's scrollport. A wrapper carrying `overflow: hidden` for an unrelated reason becomes that scroll container, so the header is pinned to the wrapper rather than to the viewport. The wrapper scrolls away with the page and takes the header with it.
blind because
Scroll offset zero is the one position at which a working and a broken sticky element are in the same place, and it is the position every capture is taken at. The computed value does not separate them either: it reads `sticky` in both cases, because the declaration won the cascade and the defect lives in an ancestor.
the check
Scroll and re-measure, rather than reading the declaration or the resolved value: `[0, 800, 2000].map(y => { scrollTo(0, y); return el.getBoundingClientRect().top; })`. Observed with Chrome 151 headless on the same markup twice: with a plain wrapper the header reported top 0, 0, 0; with `overflow: hidden` on that wrapper it reported 0, -800, -2000, having left the viewport entirely. `getComputedStyle(el).position` returned 'sticky' in both runs.
cost of missing
The navigation is unreachable on every long page, and the property is re-declared, re-prefixed and re-tested on the element while the ancestor that revoked it is never examined.
mitigation
Walk the ancestors and check for a scroll container: any ancestor whose computed overflow is not `visible` in the sticky axis is the one the element is pinned to.
generalises to
Any property whose effect is decided by an ancestor rather than by the element declaring it: sticky and fixed positioning, stacking contexts, percentage heights, transforms creating containing blocks.
A full-page capture paints a fixed element once, at the offset it was captured from
reads as
The full-page capture is 2400 pixels tall and the consent banner appears as a stripe near the top, well clear of the call to action further down. Conclusion drawn: the banner obstructs nothing.
actually
A capture beyond the viewport renders the document once; the DevTools Protocol describes captureBeyondViewport as no more than capturing the screenshot beyond the viewport. A `position: fixed` element is painted at its viewport position at that single moment, which places it at one arbitrary document offset in the resulting image. In a browser it occupies that same band of every viewport at every scroll position.
blind because
A full-page image is a picture in document space; a fixed element lives in viewport space. Flattening one onto the other destroys exactly the property that made the element worth checking, and leaves an image in which the overlay covers a small fraction of a very tall page.
the check
Ask each fixed element what share of the viewport it owns: `[...document.querySelectorAll('*')].filter(e => getComputedStyle(e).position === 'fixed').map(e => e.className + ': ' + Math.round(e.getBoundingClientRect().height) + 'px = ' + Math.round(100 * e.getBoundingClientRect().height / innerHeight) + '% of every viewport')`. Observed with Chrome 151 headless: `["banner: 180px = 39% of every viewport"]`. The banner occupied rows 277-456 of the 457-pixel viewport capture and rows 277-456 of the 2400-pixel full-page capture, which is 39% of what a visitor sees and 7% of the image reviewed.
cost of missing
A banner, chat launcher or toolbar that permanently covers the bottom third of every screen is signed off against an image in which it covers a stripe of empty page, and the obstruction is only discovered from conversion data.
mitigation
Review fixed elements from viewport captures at several scroll offsets, and use `document.elementFromPoint` at the coordinates of anything that must remain clickable.
generalises to
Every rendering that resolves a relative frame into an absolute one: fixed positioning in full-page captures, relative timestamps baked into a report, `~` expanded at write time rather than read time.
A Content-Length shorter than the body truncates the response with no error anywhere
reads as
`curl -s -o data.json -w '%{http_code}' URL` prints 200, curl exits zero, and the bytes written match the Content-Length the server declared. Conclusion drawn: the document was retrieved intact.
actually
RFC 9112 makes the declared length authoritative: if a valid Content-Length header field is present without Transfer-Encoding, its decimal value defines the expected message body length in octets. The client reads that many and stops, whatever else is on the connection. The commonest cause is a handler that computes the length in characters and writes the body in UTF-8, so the response is cut short by exactly the number of extra bytes the non-ASCII characters cost. RFC 9112 anticipates the remainder: a user agent MAY discard the remaining data or attempt to determine if that data belongs as part of the prior message body, which might be the case if the prior message's Content-Length value is incorrect.
blind because
Every length-based check agrees with itself. The header says 67, the file on disk is 67 bytes, and `%{size_download}` is 67. Unlike a connection that dies mid-body, nothing here is incomplete from the transport's point of view, so no exit code, no warning and no retry is produced.
the check
Validate the body on its own terms rather than on the sender's: `curl -s URL | python3 -c 'import sys, json; json.load(sys.stdin)'`. Observed against a local handler serving a 73-byte UTF-8 JSON document under `Content-Length: 67`, the length of the same text in characters: curl reported code=200 size_download=67 and exited 0; the saved file ended `"ok":` and json.load raised `JSONDecodeError: Expecting value: line 1 column 62`. The identical handler taking its length from the encoded bytes returned 73 and parsed. A separate server declaring 16 against a 131-byte body gave curl, Python's http.client and Node all the same silent 16-byte prefix with status 200.
cost of missing
A record set is short by a few entries, a token is cut mid-string, a page loses its closing markup. Each looks like a data problem upstream, and re-requesting reproduces it exactly, which reads as confirmation rather than as a framing bug.
mitigation
Check completeness at the semantic layer for anything whose length matters: parse it, verify its terminator, or compare a digest the origin computed over the same bytes.
generalises to
Every length or checksum supplied by the same party that produced the payload, and every count taken in one unit and spent in another.
Two field lines with one name collapse into one value, and clients disagree about which
reads as
`resp.headers['X-Frame-Options'] == 'DENY'` and the assertion passes. Conclusion drawn: the response carries the header the policy requires.
actually
The response carried the field twice, DENY then ALLOWALL, because two layers each added it. RFC 9110 permits recombination only in limited circumstances: a sender MUST NOT generate multiple field lines with the same name in a message unless that field's definition allows multiple field line values to be recombined as a comma-separated list. Senders do it anyway, and recipients then differ. Some return the first value, some the joined list, some keep both. X-Frame-Options is not a list-valued field, so a browser receiving it twice with conflicting values has no defined behaviour to fall back on.
blind because
A header dictionary maps one name to one string. It has no way to represent 'this name appeared twice with contradicting values', so the single reading that satisfies the assertion is the only reading the instrument can produce.
the check
Read the field lines rather than the parsed mapping: `curl -sD - -o /dev/null URL | grep -ci '^x-frame-options:'`, and treat any count above one as a failure. Observed against a local server sending X-Frame-Options twice (DENY, ALLOWALL) and Cache-Control twice (no-store, max-age=31536000): curl printed both lines each time; Python's urllib.request returned 'DENY' and 'no-store' from `headers[name]` while `headers.get_all` returned both values; `http.client.getheader` returned the joined 'no-store, max-age=31536000'; Node 20.20.2 returned the joined 'DENY, ALLOWALL'. One response, three clients, three different answers.
cost of missing
A security-header audit passes on a response no browser will honour, and a caching audit reads no-store while a shared cache reads the joined value and stores the response for a year.
mitigation
Assert on the count of field lines as well as on the value, and use the client API that preserves multiplicity (`get_all`, `rawHeaders`) wherever a header carries policy.
generalises to
Every flattening of a multi-valued source into a scalar: repeated query parameters, environment variables set twice, merged configuration layers, duplicate rows collapsed by a join.
NS-077 · a status codedocumentedconnection-named-differently-from-the-request
Overriding the Host header leaves the TLS handshake pointing somewhere else
reads as
`curl -H 'Host: app.example.com' https://<origin-ip>/` returns 200 with a plausible page. Conclusion drawn: the origin serves app.example.com correctly.
actually
The hostname in the URL, not the Host header, supplies the TLS server_name extension. RFC 6066: a server that receives a client hello containing the server_name extension MAY use the information contained in the extension to guide its selection of an appropriate certificate to return to the client, and/or other aspects of security policy. curl documents the split plainly for --connect-to, which is only used to establish the network connection and does NOT affect the hostname/port that is used for TLS/SSL (e.g. SNI, certificate verification). With a bare IP in the URL no SNI is sent at all, so a terminator that selects a certificate, a backend or a WAF policy by SNI falls through to its default.
blind because
The 200 is real and the body may well be the right site, because the default virtual host is often the site under test. Nothing in the status, the body or the headers records which name the handshake carried.
the check
Keep the hostname in the URL and move only the address: `curl -sk --resolve app.example.com:443:<ip> https://app.example.com/`. Observed against a local TLS server that reports both values in its body: `curl -sk -H 'Host: canary.example' https://127.0.0.1:8443/` returned `SNI=None HOST=canary.example`, while `curl -sk --resolve canary.example:8443:127.0.0.1 https://canary.example:8443/` returned `SNI=canary.example HOST=canary.example:8443`. `openssl s_client` split the same way: no -servername, no SNI.
cost of missing
A certificate, a routing rule or a security policy bound to the hostname is never exercised, so the probe passes for a name that would fail for every real client, and the failure appears only when DNS is cut over.
mitigation
Use --resolve or --connect-to, which change where the connection goes without changing what the request and the handshake claim to be.
generalises to
Every request whose identity is asserted at more than one layer: SNI against Host, DNS name against certificate subject, connection string host against database catalogue, tenant header against signed token.
The first status line in a response may belong to an interim response
reads as
`curl -sD headers.txt URL` succeeded and headers.txt begins with a status line followed by a header block. Conclusion drawn: those are the response's status and its headers.
actually
An origin sending 103 Early Hints emits a complete status line and header section before the real one. RFC 9110 requires clients to cope: a client MUST be able to parse one or more 1xx (Informational) responses received prior to a final response, and such a response terminates when the header section ends. A parser that stops at the first blank line, which is what the message grammar tells it to do, reads the interim block, whose header section commonly contains nothing but `link`.
blind because
Both blocks are well-formed HTTP with the same shape. Nothing marks the first as provisional except its status code, which is precisely the field being taken on trust, and whether the interim block appears at all depends on the protocol version negotiated rather than on the URL.
the check
Count the status lines before reading any of them: `curl -sD - -o /dev/null --http2 URL | grep -c '^HTTP/'`, and treat anything above one as two blocks to disentangle. Observed at 00:28 UTC on 2026-08-24 against https://www.cloudflare.com/: 2 under --http2, with `head -1` returning `HTTP/2 103` and `%{http_code}` returning 200; 1 under --http1.1, where the same origin sent only `HTTP/1.1 200 OK`.
cost of missing
A header audit reads the Early Hints block, finds one `link` field, and reports HSTS, CSP, Content-Type and Cache-Control as absent from a response that carries all four. A status check that reads the first line records a 103 and either alerts or, worse, treats an unknown class as a pass.
mitigation
Take the status from the client's own final-response accessor (`%{http_code}`, `response.status`) rather than from the first line of a header dump, and parse header dumps as a sequence of blocks.
generalises to
Every stream where a provisional message precedes the real one: 1xx responses, redirect chains, retried requests, partial results emitted before a final aggregate.
Field names arrive lowercased over HTTP/2, so a case-sensitive check passes vacuously
reads as
The audit greps the response headers for `^Set-Cookie:` lines lacking the Secure attribute, finds none, and records a pass. Conclusion drawn: no insecure cookie is being set.
actually
RFC 9113 requires that field names MUST be converted to lowercase when constructing an HTTP/2 message. Over HTTP/2 the field arrives as `set-cookie:` and a pattern anchored to `^Set-Cookie:` matches nothing at all, neither the compliant cookies nor the insecure ones. Which casing arrives is decided by the version negotiated for that particular request, which is a property of the connection rather than of the URL, so the same command can examine everything in one environment and nothing in the next.
blind because
An absence-based assertion cannot separate 'the condition does not occur' from 'the field was never examined'. Both return zero matches, exit the same way, and read as a clean result.
the check
Prove the pattern matches something before trusting that it matched nothing, and record the version alongside it: compare `curl -sD - -o /dev/null --http1.1 URL | grep -c '^Content-Type:'` against the same command with `--http2`. Observed at 00:28 UTC on 2026-08-24 against https://www.cloudflare.com/: 1 under --http1.1, 0 under --http2, and 1 under --http2 for `grep -c '^content-type:'`. `--http2` had also negotiated 1.1 without complaint against a local HTTP/1.1-only server, reporting `version=1.1 code=200` and exiting 0.
cost of missing
A control that has never once evaluated its subject reports a pass on every run, and the run history makes it look like a stable, long-standing green.
mitigation
Match case-insensitively, and give every absence-based check a positive control that must match for the check to be considered to have run.
generalises to
Every check whose passing condition is an empty result set: greps for a forbidden pattern, queries returning no rows, log searches finding no errors, linters with a misconfigured include path.
cd with an empty or unset argument succeeds without going anywhere
reads as
`cd "$BUILD_DIR" && rm -rf ./*` exits zero and the script proceeds. Conclusion drawn: the build directory was entered and cleared.
actually
bash(1): change the current directory to dir; if dir is not supplied, the value of the HOME shell variable is the default. An unset variable left unquoted disappears during expansion, so `cd $BUILD_DIR` becomes `cd` and succeeds by moving to the home directory. Quoted, `cd ""` also succeeds and leaves the working directory exactly where it was. Both return zero and print nothing, and the destructive command that follows runs wherever the shell happened to be.
blind because
cd's status reports whether a directory change was performed, never which directory. Arriving where you intended and arriving in the home directory are the same value.
the check
Confirm the destination rather than the status, or refuse an empty value outright: `: "${BUILD_DIR:?BUILD_DIR is empty}"; cd "$BUILD_DIR" && [ "$PWD" = "$BUILD_DIR" ]`. Observed on bash 5.2.21 from a scratch directory: with TARGET unset, `cd $TARGET` exited 0 and left PWD at the user's home directory; with TARGET set to the empty string, `cd "$TARGET"` exited 0 and left PWD unchanged. `set -u` caught only the unquoted unset case, and `set -eu` ran straight past the quoted empty one with status 0. With CDPATH=/usr, `cd bin` from /tmp exited 0 in /usr/bin.
cost of missing
The rm, rsync or build step that follows operates on the wrong tree with full confidence, and in the home-directory case on a tree that contains everything.
mitigation
`${VAR:?}` fails on empty as well as unset, which `set -u` does not; assert on $PWD after any cd whose argument came from a variable.
generalises to
Every command that treats a missing argument as a request for its default rather than as an error: cd, `git checkout`, `kubectl` without a namespace, `docker build` with an empty context path.
A declaration builtin consumes the exit status of the substitution it assigns
reads as
`local token=$(fetch_token)` is followed by a status check, and the check passes. Conclusion drawn: fetch_token succeeded and token holds a token.
actually
bash(1) explains the ordinary case: if no command name results and one of the expansions contained a command substitution, the exit status of the command is the exit status of the last command substitution performed. Putting `local`, `declare`, `export` or `readonly` in front supplies a command name, so the status becomes that builtin's instead, and the return status is 0 unless local is used outside a function, an invalid name is supplied, or name is a readonly variable. The substitution's failure is discarded and the variable holds an empty string.
blind because
One line performs two operations and reports on the outer one. The status is a true statement about whether a variable was declared, offered where a statement about whether a value was obtained is expected.
the check
Separate the declaration from the assignment and compare the two forms: `local token; token=$(fetch_token)`. Observed on bash 5.2.21: `f(){ local out; out=$(false); echo $?; }` printed 1, `g(){ local out=$(false); echo $?; }` printed 0, and `h(){ export OUT=$(false); echo $?; }` printed 0. Under `set -e` the split form aborted the shell and the combined form ran on to completion returning 0.
cost of missing
An empty credential, empty version string or empty path is carried forward and fails somewhere far from its origin, usually as an authentication error or a path that resolves to the filesystem root.
mitigation
Declare on one line and assign on the next wherever the command's outcome matters; `set -e` only helps once the two are separated.
generalises to
Every wrapper that reports on itself rather than on what it invoked: shell builtins in front of assignments, test harnesses swallowing setup failures, entrypoints exiting on the shell's status rather than the program's.
An unmatched pattern is passed through as a literal filename
reads as
`for f in releases/*.tar.gz; do verify "$f"; done` exits zero and the script reports the release set verified. Conclusion drawn: every archive was checked.
actually
bash(1): if no matching filenames are found, and the shell option nullglob is not enabled, the word is left unchanged. The loop therefore runs exactly once, with the variable set to the literal string `releases/*.tar.gz`, which names nothing. Any command that tolerates a missing operand, `rm -f` and `mkdir -p` and `grep -s` among them, returns zero, and so does the loop.
blind because
Zero iterations and one iteration over a path that does not exist yield the same exit status. The difference lives in a count that nothing reports, and a wrong directory, a typo in an extension and an empty build all produce it identically.
the check
Count what the pattern matched instead of what the loop returned: `shopt -s nullglob; files=(releases/*.tar.gz); echo "${#files[@]}"`, and fail on zero. Observed on bash 5.2.21 in an empty directory: `for fn in *.log; do echo "[$fn]"; done` printed `[*.log]` and exited 0; `rm -f *.log` exited 0 having deleted nothing; the same loop under `shopt -s nullglob` ran zero iterations.
cost of missing
A cleanup, upload or signing step reports success over an empty set, so the artefacts it was meant to handle survive untouched, and the next stage consumes the previous release without noticing it is stale.
mitigation
Enable `nullglob` and assert on the array length, or `failglob` where an empty match is always a bug.
generalises to
Every operation over a collection that is silently empty: globs, empty pipelines into xargs, queries returning no rows, iterations over an unset list.
An in-place edit of a symlink replaces the link with a regular file
reads as
`sed -i 's/old/new/' app.conf` exits zero and reading app.conf shows the new value. Conclusion drawn: the configuration was updated.
actually
GNU sed edits in place by writing a temporary file and renaming it over the target, and it does not resolve symbolic links unless asked; the existence of `--follow-symlinks`, documented as following symlinks when processing in place, is the acknowledgement. The rename replaces the link itself. The path now holds a regular file carrying the new content, the file the link pointed at is untouched, and the link no longer exists.
blind because
Reading the path returns the edited content, because the path genuinely holds it now. What changed is the identity of the object behind the name, and content is the one thing that cannot reveal it.
the check
Look at the type and inode behind the name rather than at the bytes: `stat -c '%i %F %N' app.conf`. Observed on GNU sed 4.9 with app.conf a symlink to repo/app.conf: beforehand `2659981 symbolic link 'app.conf' -> 'repo/app.conf'`; after `sed -i`, `2659983 regular file 'app.conf'` holding `setting=new`, while repo/app.conf still held `setting=old` at its original inode 2659980. With `--follow-symlinks` the link survived and repo/app.conf received the edit.
cost of missing
The change is invisible to the repository the link came from, is not committed, and is silently reverted the next time the link farm is rebuilt or the host is reprovisioned. Meanwhile every other path sharing the original inode still serves the old value.
mitigation
`sed -i --follow-symlinks`, or edit the resolved path from `readlink -f`. The same applies to any tool that writes by rename.
generalises to
Every write-by-rename: editors saving atomically, `sed -i`, `sort -o`, dotfile farms, hard links and bind mounts that expected to keep sharing an inode.
A file that changed without changing size or timestamp is never transferred
reads as
`rsync -a src/ dst/` exits zero, and dst/app.conf exists with the same size and the same modification time as the source. Conclusion drawn: the destination is a copy of the source.
actually
rsync(1): rsync finds files that need to be transferred using a 'quick check' algorithm (by default) that looks for files that have changed in size or in last-modified time. A file edited in place to the same length, restored from an archive that preserved timestamps, or written by a generator that copies mtime from its input matches on both counts and is skipped. The old content remains, and the transfer is reported as complete because nothing needed transferring.
blind because
The two numbers the tool decides on are the two numbers used to verify it. Size and mtime agree, which is a true statement about the metadata and a false one about the contents.
the check
Compare contents rather than metadata: `rsync -ain --checksum src/ dst/` lists exactly what a content comparison would move. Observed on rsync 3.2.7 with src/app.conf holding VERSION=2 and dst/app.conf holding VERSION=1, both 10 bytes with mtime forced to 2026-01-01: `rsync -av src/ dst/` exited 0, reported `sent 72 bytes`, listed no files, and left the destination at VERSION=1. The same pair under `--checksum` transferred and the destination became VERSION=2. `cp -u` copied nothing for the same reason.
cost of missing
A deploy or backup succeeds and changes nothing, and repeating it reproduces the same success, so the obvious response to the symptom confirms the wrong hypothesis.
mitigation
Use `--checksum` for any sync where the content is authoritative and the timestamps are not, accepting the read cost; or ensure the writer touches mtime whenever it rewrites.
generalises to
Every change detector keyed on a proxy for content: mtime-based build systems, ETags derived from metadata, cache keys built from a version string that was not bumped.
A repeated key is resolved silently, and the occurrence you read is not the one in force
reads as
config.json declares `"debug": false` and a database host of prod.db.internal, and it is the file the service loads. Conclusion drawn: debug is off and the service talks to production.
actually
The same names appear again further down the file, added by a later edit or a careless merge. RFC 8259: the names within an object SHOULD be unique, and when the names within an object are not unique, the behavior of software that receives such an object is unpredictable; many implementations report the last name/value pair only. The values in force are the ones at the bottom.
blind because
Reading a configuration file means reading downwards and stopping at the first occurrence of the key in question. Nothing in the syntax marks a name as later overridden, and both occurrences are individually valid.
the check
Load it through the parser the service uses and print what it produced: `python3 -c 'import json, sys; print(json.load(open(sys.argv[1])))' config.json`. Observed on Python 3.12.3, Node 20.20.2 and jq against a file declaring debug false then debug true and a database host prod.db.internal then localhost: all three produced `{'debug': True, 'database': {'host': 'localhost'}}`, and `jq keys` reported two keys rather than four. PyYAML 6.0.1 behaved the same way on the YAML equivalent. Python's configparser instead raised DuplicateOptionError, so whether the file is accepted at all depends on which parser reads it.
cost of missing
Debug output, a staging database or a permissive CORS origin is live in production, and the file that proves otherwise is the same file that enables it.
mitigation
Reject duplicates at load time; Python's `json.load` accepts an `object_pairs_hook` that can raise on a repeated name, and most YAML loaders can be configured to do the same.
generalises to
Every last-writer-wins merge that leaves both writers visible: duplicate keys, repeated environment assignments, layered configuration, stylesheet rules of equal specificity.
kill reports success when the signal was delivered and disregarded
reads as
`kill $PID` exits zero and the deploy script moves on. Conclusion drawn: the old worker has stopped.
actually
kill(2): on success, at least one signal was sent, zero is returned. Success means the signal was queued to a process the caller had permission to signal. A process that installed an ignore disposition for SIGTERM, whether through `trap '' TERM` or a runtime that swallows it while a shutdown hook stalls, receives the signal and carries on. signal(7) notes the only exceptions: SIGKILL and SIGSTOP cannot be caught, blocked, or ignored.
blind because
The status describes the sender's half of the transaction. Whether the recipient acted is a fact about the recipient, observable only afterwards and only by looking again.
the check
Read the target's signal dispositions, or simply look again after a pause: `grep -E '^Sig(Ign|Blk|Cgt)' /proc/$PID/status`. Observed on Linux 6.8 with a script carrying `trap '' TERM` and `trap '' HUP`: two successive `kill` invocations both exited 0 and the process was still listed by `ps` after each, reporting `SigIgn: 0000000000004005`, the bits for signals 1, 3 and 15. `kill -9` ended it. `os.kill` against an unreaped zombie likewise raised nothing and returned normally.
cost of missing
The deploy continues believing the port is free. The replacement either fails to bind, or binds elsewhere and serves alongside the process that was supposed to be gone, producing a fleet where half the requests run old code.
mitigation
Treat termination as a condition to be waited on rather than an instruction to be issued: poll for the process to disappear, with a bounded escalation to SIGKILL.
generalises to
Every asynchronous request whose acknowledgement is acceptance of the message rather than performance of the work: signals, queue publishes, webhook deliveries, cache invalidations.
A unit reported inactive can still have every worker it started running
reads as
`systemctl stop app` exits zero and `systemctl is-active app` prints inactive. Conclusion drawn: the service and everything it spawned are stopped.
actually
With KillMode=process, systemd.kill(5) states that only the main process itself is killed (not recommended!), and warns that this allows processes to escape the service manager's lifecycle and resource management, and to remain running even while their service is considered stopped and is assumed to not consume any resources. The workers keep their sockets, locks and memory. The default, control-group, kills the whole cgroup and does not have this behaviour.
blind because
is-active reports the unit's state, and the unit's state is decided by its main process. Once the survivors have outlived the unit they are no longer accounted to it, so the supervisor's view is accurate and incomplete at once.
the check
Ask the kernel who is alive rather than asking systemd whether the unit is: `ps -eo pid,ppid,args | grep '[w]orker'`, or `ss -ltnp` for the port the service held. Observed on systemd 255 with a user unit `Type=simple` and `KillMode=process` whose ExecStart backgrounded a child: `systemctl --user stop` exited 0, `is-active` printed inactive, and `ps` still listed the child at pid 3917771. The identical unit at the default KillMode=control-group left nothing behind.
cost of missing
A restart appears to work while the previous generation continues serving; two versions run concurrently and diverge, and the resources the orphans hold are invisible to anything that accounts by unit.
mitigation
Leave KillMode at control-group unless there is a specific reason not to, and verify a stop by the absence of processes and listeners rather than by the unit's state.
generalises to
Any supervisor whose notion of the service is narrower than the set of processes the service created: init systems, container runtimes, CI job runners, test harnesses spawning fixtures.
Records longer than a pipe's atomic limit are spliced into one another
reads as
`grep 'request_id=abc123' app.log` returns nothing, and the file around that period is full of well-formed lines. Conclusion drawn: that request never reached this service.
actually
pipe(7): POSIX.1 says that writes of less than PIPE_BUF bytes must be atomic, the output data being written to the pipe as a contiguous sequence, while writes of more than PIPE_BUF bytes may be nonatomic, and the kernel may interleave the data with data written by other processes. On Linux PIPE_BUF is 4096 bytes. Several workers writing to one pipe, which is what a container's stdout, a `tee` and most log shippers are, produce records cut open with another worker's record inserted into the gap. The result still ends in a newline, so it is still a line.
blind because
A log reader sees lines, and a spliced line is a line: it has a beginning, an end and plausible contents. The pattern that would have matched now straddles a boundary that did not exist when the record was written, and the line count is unchanged.
the check
Validate each line against the format the writer emits and count the failures, rather than counting lines. Observed on Linux 6.8 with four writers into one pipe behind a deliberately slow reader: at 4090-byte records, 800 lines and 0 malformed; at 5000-byte records, 800 lines and 21 malformed; at 20000-byte records, 800 lines and 195 malformed, one of which opened with `BEGIN-B-0010` and contained an entire `BEGIN-A-0000 ... END-A-0000` record inside it. The line count was 800 in every run.
cost of missing
Requests appear never to have happened, error rates read low, and the records that would contradict both are present in the file in a form no query will match.
mitigation
Keep each record under PIPE_BUF, or give each writer its own descriptor opened O_APPEND onto a regular file, where appends do not interleave regardless of size.
generalises to
Any shared append-only channel with an atomicity limit: pipes, datagram sockets, unlocked file writes, records assembled from several write calls.
A log search returns nothing because the window and the timestamps are in different zones
reads as
`journalctl -u app --since '2026-08-24 00:20:00'` prints `-- No entries --` for a window that covers the incident. Conclusion drawn: the service logged nothing then, so it was not running or was never reached.
actually
journalctl interprets --since and --until in the local time zone, and systemd.time(7) states that on display systemd will format timestamps in the local timezone. When the window is copied from a source in another zone, a UTC dashboard, a cloud console, an API response or a colleague on another continent, the query addresses a moment hours away from the one intended. The entries exist and sit outside the range.
blind because
An empty result set has one shape. Nothing separates 'no entries in this window' from 'the window was somewhere else', and the timestamps that would reveal the offset are precisely the ones the filter excluded.
the check
Ask for the entries in an unambiguous frame and see whether they exist at all before filtering: `journalctl -u app -n 5 --utc -o short-iso`. Observed on this box (Etc/UTC) against a single `logger -t vftz` entry: plain `journalctl -t vftz` displayed it as `Aug 24 00:24:27`, while `TZ=America/New_York journalctl -t vftz` displayed the same entry as `Aug 23 20:24:27`, a different calendar day. Passing a window taken from the UTC clock while TZ was America/New_York returned `-- No entries --` for a record written seconds earlier.
cost of missing
The investigation concludes the service was silent during the incident and moves upstream, while the evidence sits in the same file a few hours away.
mitigation
Pin both ends of every correlation to one frame: query with --utc and read with -o short-iso, or attach an explicit offset to every timestamp that crosses a system boundary.
generalises to
Every filter expressed in units the store does not share: time zones, seconds against milliseconds, inclusive against exclusive bounds, severities named differently by the writer and the query.
Crawler hits in an access log are not evidence that a page is indexed
reads as
The access log shows repeated fetches from Googlebot, Bingbot and other declared crawlers, and the sitemap was accepted. Conclusion drawn: the pages are in the index and the site is discoverable.
actually
Crawling, indexing and ranking are three separate stages. A crawler fetching a URL records only that it was retrieved; the page may then be excluded, deduplicated against similar content, or held in a queue for days. A site can be fetched hundreds of times and return no results for a search of its own exact title.
blind because
The access log is written by the origin and can only record requests that reached it. It has no field for what the requester did afterwards, and no stage of indexing produces a request back to the server.
the check
Query the index itself rather than reading the log: search for an exact phrase unique to the page, in quotes, and separately run a `site:` query for the domain. Both return nothing while the page is merely crawled. For a property you control, the index-coverage report in Google Search Console or Bing Webmaster Tools states the stage per URL.
cost of missing
Distribution work is reported as finished on the strength of crawler traffic, and the weeks in which the pages are fetched but unfindable pass unnoticed. Effort moves on to new content while nothing published so far can be reached by search.
mitigation
Treat submission and crawling as inputs, not outcomes. The observable outcome is a result page containing your URL.
generalises to
Any pipeline whose early stages report back to you and whose later stages do not: submitted-versus-accepted, queued-versus-delivered, uploaded-versus-published, deployed-versus-serving.