VERIFYFIRST — what your verification method cannot see
https://verifyfirst.dev
A failure that reports success costs more than one that crashes. This is a reference for the moment before you claim work is done: you are about to verify through some instrument, and every instrument is structurally blind to something. Look up what yours cannot see.
Entries are organised by the instrument that missed the failure, not by the technology involved. Each states the false reading, the true state, why the instrument cannot separate them, and one discriminating check. A check qualifies only if it returns different output under the two hypotheses.
Every entry is a failure that genuinely occurs. Entries marked provenance 'observed' were diagnosed first-hand during the work that produced this site. Entries marked 'documented' cite primary documentation and their discriminating check was reproduced before publication. None are hypothetical.
90 entries / 6 instruments / 31 symptoms / v7.2.0 / 2026-08-24
CC0-1.0. Everything below is public domain.
======================================================================
SYMPTOMS — what you are seeing
======================================================================
* The page renders blank
Nothing visible, but the DOM may be complete and merely invisible.
-> NS-010, NS-002, NS-004, NS-052, NS-072
* The deploy ran but nothing changed
The new code is on disk. Something between disk and user is still serving the old one.
-> NS-005, NS-007, NS-008, NS-026, NS-038, NS-058, NS-059, NS-084
* The command succeeded but had no effect
Exit zero describes the process, not the outcome you wanted.
-> NS-005, NS-015, NS-016, NS-014, NS-053, NS-055, NS-056, NS-080, NS-081, NS-082, NS-086
* The service says active but is not working
Active is a statement about a process existing, not about it serving.
-> NS-026, NS-027, NS-005, NS-037, NS-038, NS-087
* An animation or counter never moves
Frozen at frame zero is indistinguishable from correctly static.
-> NS-001, NS-002
* The tests pass but the feature is broken
A test that never ran, and a test that asserts nothing, both report green.
-> NS-017, NS-018
* The API returned 200 but the data is wrong
The status describes the transaction, not the payload.
-> NS-020, NS-021, NS-019, NS-025, NS-042, NS-045, NS-057, NS-058, NS-075, NS-076, NS-078
* My config change is being ignored
The file records intent. Something later, or something else, decided the outcome.
-> NS-003, NS-023, NS-008, NS-046, NS-060, NS-061, NS-062, NS-083, NS-084, NS-085
* The logs show nothing useful
Silence is produced by a filter, a rotation and a crash alike.
-> NS-029, NS-028, NS-011, NS-031, NS-032, NS-033, NS-034, NS-066, NS-067, NS-068, NS-088, NS-089
* A process died and I cannot tell why
The thing that was killed is often not the thing that caused it.
-> NS-011, NS-013, NS-014, NS-031
* A URL returns 200 for something that does not exist
A catch-all answered on the origin's behalf.
-> NS-022, NS-007, NS-043
* Clicks land on nothing
The element you can see is not the element receiving the event.
-> NS-030
* The wrong file or image was used
A plausible artifact of the right type is not evidence it is the right one.
-> NS-006, NS-004, NS-048, NS-062, NS-056, NS-083, NS-077
* Text renders, but it looks wrong
A substitution that still renders is invisible without a comparison.
-> NS-004, NS-024, NS-049, NS-051
* Assets take seconds to appear
Late departure and slow transfer feel identical from the viewport.
-> NS-009
* Numbers come back subtly altered
Silent coercion produces a valid value that is not your value.
-> NS-025, NS-019, NS-046
* A search for a process finds something unexpected
The instrument is a process, and it is inside its own sample.
-> NS-012, NS-027, NS-036, NS-039, NS-040, NS-041, NS-086, NS-087
* A command hangs and never returns
There is no exit code to read, and silence resembles progress.
-> NS-014, NS-013
* Log lines are in an impossible order
Two streams with different buffering do not interleave the way they were written.
-> NS-035, NS-066, NS-088
* Credentials are ignored but the request still succeeds
Something in the chain dropped the header and the endpoint answered anyway.
-> NS-044
* A fresh clone is missing files that are present locally
Committed is not the same as tracked.
-> NS-047, NS-060
* The screenshot does not match what I see in a browser
A capture is taken under its own device scale, colour scheme and stylesheet, not yours.
-> NS-049, NS-050, NS-051, NS-052, NS-070, NS-071, NS-072, NS-073, NS-074
* A downloaded file is not the content I expected
The transfer succeeded; the encoding or the destination path did not.
-> NS-057, NS-054, NS-056, NS-075
* The machine looks busy but nothing is progressing
Load and CPU percentages measure different things than they appear to.
-> NS-064, NS-065, NS-063
* A file I edited lost its contents
The shell truncates a redirect target before the command reading it ever starts.
-> NS-054, NS-062
* The process ran out of memory but the box has plenty
A container sees the host's totals, not its own limit.
-> NS-063
* My analytics and my server logs disagree
One of them is counting what clients claim to be, and clients set that field themselves.
-> NS-069, NS-090
* A shell variable was empty and nothing complained
Several builtins treat an empty argument as a request to do nothing, successfully.
-> NS-080, NS-081, NS-082
* curl and my HTTP library disagree about the same response
Header collapsing, protocol version and informational status lines differ per client.
-> NS-076, NS-078, NS-079
* grep finds nothing in a header dump that clearly contains it
HTTP/2 lowercases every field name; an anchored pattern matches neither case.
-> NS-079
* Nobody can find the thing I published
Being fetched, being indexed and being findable are three different states.
-> NS-090
======================================================================
INSTRUMENT: screenshot — A rendered image
======================================================================
Used when: You captured the page and looked at it.
Captures: One frame's worth of pixels, at one viewport size, at one moment.
Cannot see:
- Time. A still frame cannot distinguish 'renders one frame then stops' from 'renders one frame correctly'.
- Whether scripts ran at all, as opposed to running and producing this.
- Why an element is absent: never drawn, drawn transparent, drawn offscreen, or covered.
- Which typeface actually resolved, when the fallback is also a real font.
- Whether an asset arrived slowly or was requested late.
NS-001 The compositor-free browser reports frozen animation as no animation
reads as A screenshot shows the element static. Conclusion drawn: the animation is badly designed, or the values are wrong.
actually requestAnimationFrame never fires because the headless browser has no compositor. Every rAF-driven counter, canvas loop and scroll handler is frozen at frame zero, whatever its quality.
blind because A still image cannot distinguish 'renders one frame then stops' from 'renders one frame correctly'. Both produce the same pixels.
CHECK let n=0; requestAnimationFrame(()=>n++); setTimeout(()=>console.log('rAF fired:', n), 1000)
cost Every visual parameter gets tuned against a frame the animation never advances past. The tuning is not merely useless, it is fitted to an artifact.
generalises Any observation instrument that shares a failure mode with the thing observed.
provenance observed
NS-002 A var assignment silently overwrites a hoisted function of the same name
reads as A canvas is blank and unanimated. Conclusion drawn: a rendering or design problem.
actually The file already used `var start` as a timestamp. A later `function start()` was hoisted, then overwritten by the number at execution. The call that scheduled initialisation threw a TypeError and init never ran. The element sat at its untouched default size with zero painted pixels.
blind because An element that renders nothing and an element that renders badly both look like a design problem in a screenshot. Nothing distinguishes them visually.
CHECK Read the element's backing store, not its appearance: canvas.width/height still at the 300x150 default means resize() never ran.
cost Four rounds of visual tuning applied to a layer that was never drawing.
generalises Any name reused across a value and a declaration in the same scope.
provenance observed
NS-004 A relative font URL resolves one directory too deep and fails into a plausible fallback
reads as Text renders in a serif. Conclusion drawn: the font loaded.
actually A stylesheet at /assets/fonts.css requesting url('assets/fonts/x.woff2') resolves to /assets/assets/fonts/x.woff2 and 404s. The browser substitutes a system serif without complaint.
blind because The fallback is a working font. Only someone who knows the intended typeface can see the substitution, and only by comparison.
CHECK document.fonts.check('1em "Family Name"') or a 404 on the font path in the network log.
cost Design review proceeds against the wrong typeface. Every judgement about weight, rhythm and scale is made on a substitute.
generalises Every fallback that is good enough to pass inspection: default configs, cached credentials, stub implementations.
provenance observed
NS-009 Assets are not slow, they are queued behind synchronous work
reads as Images take seconds to appear. Conclusion drawn: the images are too large.
actually Their request had not been issued. Heavy synchronous work earlier in the document held the main thread, and the fetch that would load them sat unsent behind it.
blind because Slow arrival and late departure are the same experience from the viewport.
CHECK performance.getEntriesByType('resource') — startTime separates 'requested late' from 'transferred slowly'.
cost Assets get compressed, resized and lazy-loaded, which lowers quality without touching the delay.
generalises Any queue where wait time is read as service time.
provenance observed
NS-010 Reveal-on-scroll renders a blank page when the observer never fires
reads as A page loads blank in an embedded or scripted context. Conclusion drawn: a rendering failure.
actually Elements start at opacity 0 and are revealed by an IntersectionObserver callback. Where the observer does not fire, the page is fully present and fully invisible.
blind because The DOM is complete and correct. Only computed opacity distinguishes it from a page that failed to build.
CHECK Compare element count against visible count: document.querySelectorAll('.reveal').length versus those with computed opacity above zero.
cost A working page is diagnosed as broken. Worse, the reverse: a genuinely blank page is dismissed as this.
generalises Every design where the default state is invisible and visibility depends on a callback.
mitigation Any progressive-enhancement pattern that hides content by default needs a timeout that shows it regardless.
provenance observed
NS-030 A transparent overlay takes the click the screenshot shows landing on the button
reads as The screenshot shows the button unobscured and correctly placed, and the click was dispatched without error. Conclusion drawn: the button was clicked.
actually A transparent element — a full-viewport modal backdrop, a zero-opacity loading layer, an oversized decorative pseudo-element — covers the button's centre point, and hit testing delivers the event to the topmost element at that coordinate. WebDriver has a named error for exactly this: the Element Click command could not be completed because the element receiving the events is obscuring the element that was requested clicked.
blind because A transparent overlay contributes no pixels. The image of a covered button and the image of an uncovered one are the same image.
CHECK Ask the document what occupies the point: `const r = el.getBoundingClientRect(); document.elementFromPoint(r.left + r.width/2, r.top + r.height/2) === el` — true when the element would receive the click, false when something is over it.
cost An automated flow reports submitting forms it never submitted. A synthetic `el.click()` compounds it, because dispatching on the element directly bypasses hit testing and succeeds where a real user's click would not.
generalises Any interaction verified by appearance rather than by the effect the interaction was supposed to have.
source https://www.w3.org/TR/webdriver2/#errors
provenance documented
NS-049 A screenshot's pixel grid is not the page's coordinate grid
reads as The capture shows the button with its centre at image pixel (400, 140), and a click is dispatched at (400, 140). Conclusion drawn: the click landed on the button.
actually The capture was taken at a device scale factor above one, so image pixels and CSS pixels differ by that factor. devicePixelRatio is the ratio between the size of a device pixel and the size of a CSS pixel, and a screenshot is measured in the former while every scripting and automation coordinate is expressed in the latter. At scale 2 the button's centre is at CSS (200, 70); (400, 140) is a different part of the page.
blind because The image carries no units. A 1600-pixel-wide PNG of an 800-pixel-wide viewport and a 1600-pixel-wide PNG of a 1600-pixel-wide viewport are both simply wide images.
CHECK Compare the capture's pixel dimensions against the page's own report of its viewport: `window.innerWidth` and `window.devicePixelRatio`. Observed with Chrome 151 headless on one 800x600 window: at --force-device-scale-factor=1 the PNG was 800x600 with devicePixelRatio 1; at 2 it was 1600x1200 with devicePixelRatio 2; at 3 it was 2400x1800. The page reported an 800 CSS-pixel viewport in all three.
cost Every coordinate derived from the image is wrong by a constant factor, and the clicks land on whatever occupies the scaled position. Because something usually does, the run continues and reports the steps it believed it took.
generalises Any measurement taken in one unit system and spent in another: viewport against document coordinates, physical against logical resolution, bytes against characters.
mitigation Take coordinates from the DOM via getBoundingClientRect, or divide image coordinates by the scale factor the capture was made at.
source https://www.w3.org/TR/cssom-view-1/#dom-window-devicepixelratio
provenance documented
NS-050 A PDF capture renders the print stylesheet rather than the page under review
reads as The page was captured to PDF and the PDF is legible and complete. Conclusion drawn: this is what the page looks like.
actually PDF generation switches the media type. Puppeteer states it directly: page.pdf() 'Generates a PDF of the page with the print CSS media type', and 'To generate a PDF with the screen media type, call page.emulateMediaType('screen') before calling page.pdf()'. Every @media print rule applies and every @media screen rule does not, so navigation, sticky headers and interactive affordances are commonly stripped by design.
blind because A PDF and a screenshot are both pictures of a page, and neither records which media type produced it.
CHECK Extract the text the capture actually contains and compare it against the screen render. Observed with Chrome 151 headless on a page carrying visible text in a .screen-only element plus `@media print{.screen-only{display:none} body::after{content:'PRINT STYLES ACTIVE'}}`: --print-to-pdf produced a file whose only text-showing operators decoded to 'PRINT STYLES ACTIVE'. The words 'Screen layout' appear nowhere in it.
cost A layout is signed off against an artifact no visitor will ever see, and print rules written long beforehand, often to remove exactly the elements being reviewed, silently define the record.
generalises Any render whose conditions are chosen by the renderer rather than the document: print media, forced colours, reduced motion, emulated devices.
source https://pptr.dev/api/puppeteer.page.pdf
provenance documented
NS-051 A headless capture exercises one branch of a colour-scheme fork
reads as The screenshot shows the page correctly styled and legible throughout. Conclusion drawn: the page renders correctly.
actually prefers-color-scheme resolves to a single value per render, and a headless browser with no desktop session reports light. Every rule inside `@media (prefers-color-scheme: dark)` was parsed, matched nothing and contributed no pixels. The dark render, which a large share of visitors receive, was never produced at all.
blind because A screenshot is one render under one set of resolved media features. The branch that did not match leaves no trace in the image, so 'the dark theme is correct' and 'the dark theme was never evaluated' look identical.
CHECK Ask the page which branch it is in, and capture both: `matchMedia('(prefers-color-scheme: dark)').matches`. Observed with Chrome 151 headless: the default run reported dark=false, light=true; the same page under --force-dark-mode reported dark=true, light=false. The page also reported prefers-reduced-motion and forced-colors as inactive by default, so those branches are unrendered for the same reason.
cost Contrast failures, unreadable text and unstyled surfaces ship in the branch nobody rendered, and remain invisible to every subsequent screenshot taken the same way.
generalises Every conditional whose condition is supplied by the environment: feature flags defaulting off, locale-dependent formatting, reduced-motion and forced-colours branches.
source https://drafts.csswg.org/mediaqueries-5/#prefers-color-scheme
provenance documented
NS-052 A capture taken at the load event shows the designed empty state
reads as The screenshot shows a clean, well-styled page reading 'No results'. Conclusion drawn: the query returned nothing, so the filter or the data is wrong.
actually The load event fires once the document and its declared subresources have loaded. It says nothing about fetches started by scripts. The request was still in flight, so the placeholder provided for a genuinely empty result was the thing on screen.
blind because The empty state is a real, intentional, correctly styled view. A picture of 'no data yet' and a picture of 'no data at all' are the same picture, because the same markup produced both.
CHECK Count the data-bearing elements at capture time instead of judging the image: `document.querySelectorAll('#list li').length`. Observed against a local endpoint delayed by two seconds: the load-event capture reported rows=0 with the empty state visible, while a capture taken after the fetch resolved reported data-rows=2 and contained `
alpha` and `beta`. The two PNGs differed in 1,262 pixels.
cost Investigation moves to the query, the filter and the backend, none of which are broken. The moment the picture was taken is the one variable never questioned.
generalises Every designed representation of absence: empty tables, zero counts, blank dashboards, 'no alerts' panels.
mitigation Trigger the capture on an assertion about content rather than on load; a network-idle condition is weaker but still better than the load event.
source https://html.spec.whatwg.org/multipage/parsing.html#the-end
provenance documented
NS-070 A capture is sized to the document, so horizontal overflow has nowhere to show
reads as The full-page capture shows every section filling the frame, with nothing clipped at either edge and no scrollbar anywhere in the image. Conclusion drawn: the layout fits the viewport.
actually Viewport-percentage units ignore scrollbars. CSS Values 4 is explicit: the viewport-percentage lengths are sized assuming that scrollbars do not exist, even if this diverges from the initial containing block. On a window 800 CSS pixels wide with a classic 15-pixel scrollbar, percentage widths resolve against 785 while 100vw resolves to 800, so every full-bleed element overhangs the layout by exactly one scrollbar and the document acquires a horizontal scrollbar of its own.
blind because A screenshot has no scrollbars and no edges beyond the content. The capture is made as wide as the document's scroll width, so the overflowing strip is inside the image rather than past its edge, and the image is the same image it would be if nothing overflowed.
CHECK Ask the document whether it is wider than its own viewport: `document.documentElement.scrollWidth - document.documentElement.clientWidth`. Observed with Chrome 151 headless at --window-size=800,600 on a vertically overflowing page: innerWidth 800, clientWidth 785, a `width:100vw` box measured 800px, a `width:100%` box measured 785px, and the difference came back as 15. On the same browser a page without any 100vw element produced a 785-pixel-wide full capture; the page with one produced an 800-pixel-wide capture, neither showing a clipped edge.
cost Every desktop visitor gets a horizontal scrollbar on every page, and on touch devices the page rubber-bands sideways. The defect is reported by users and cannot be reproduced from any capture, which sends the investigation to the wrong layer.
generalises Any measurement taken in a frame that excludes the thing being measured: viewport units against layout width, container queries against a resized container, timings taken inside the operation being timed.
mitigation Size full-bleed elements against the containing block rather than the viewport, or reserve the gutter with `scrollbar-gutter: stable` so the two frames of reference agree.
source https://www.w3.org/TR/css-values-4/#viewport-relative-lengths
provenance documented
NS-071 A full-page capture ends where the renderer decided to stop rendering
reads as The full-page capture is 2406 pixels tall, every section in it is drawn, and the article reads through to its end. Conclusion drawn: this is the whole page.
actually The sections carry `content-visibility: auto`, which CSS Containment 2 defines as turning on layout, style and paint containment and, if the element is not relevant to the user, also skipping its contents. Skipped contents are not painted, as if they had visibility: hidden, and the element is sized by its `contain-intrinsic-size` placeholder instead of by what is inside it. Offscreen sections therefore contribute their placeholder height to the document and nothing to the image.
blind because A capture records what was painted. A subtree that was never laid out contributes no pixels and no height, so a short document and a truncated one are the same picture: complete, continuous and ending in a plausible place.
CHECK Force the skipping off and re-measure the document: `(() => { const a = document.documentElement.scrollHeight; document.querySelectorAll('*').forEach(e => e.style.contentVisibility = 'visible'); return [a, document.documentElement.scrollHeight]; })()`. Observed with Chrome 151 headless on a page with three `content-visibility: auto` sections declaring `contain-intrinsic-size: auto 300px` around 718 pixels of real content each: [2406, 3654]. Each section measured 302px rather than 718px, and 1248 pixels of article were absent from the layout and from the capture alike.
cost Content review, screenshot diffing and visual regression all operate on a document that is missing most of itself, and the missing part is the part nobody scrolled to, which is where unreviewed content accumulates.
generalises Every optimisation that does less work when nobody is watching: lazy loading, virtualised lists, deferred hydration, sampled tracing.
mitigation Set `contain-intrinsic-size` to a value close to the real height so the placeholder does not distort the document, and disable content-visibility for any automated capture.
source https://www.w3.org/TR/css-contain-2/#content-visibility
provenance documented
NS-072 A frame the embedded site refused renders as ordinary whitespace
reads as The capture shows the dashboard with a clean empty band where the third-party widget sits, and the DOM confirms the iframe is present with the right src. Conclusion drawn: the widget loaded and has nothing to display.
actually The embedded document declined to be framed. RFC 7034 on X-Frame-Options: DENY means a browser receiving content with this header field MUST NOT display this content in any frame. The iframe element is still laid out at its declared size and left empty. Nothing about the parent document changes, and no layout shift marks the refusal.
blind because An iframe reserves its box before it has any content, so a frame that was refused and a frame that rendered a blank empty state occupy the same rectangle of the same colour. Inspecting the DOM confirms the element and its src, both of which are correct; what failed is on the other side of the boundary.
CHECK Ask the resource timeline what arrived rather than the DOM what exists: `performance.getEntriesByType('resource').filter(e => e.initiatorType === 'iframe').map(e => e.name + ' ' + e.transferSize)`. Observed with Chrome 151 headless against a local server: the refused frame reported transferSize 0 and logged `Refused to display 'http://127.0.0.1:8936/' in a frame because it set 'X-Frame-Options' to 'deny'`; the identical page pointed at an unprotected copy reported transferSize 405. `frames.length` was 1 and the iframe's src was the intended URL in both runs, and counting pixels inside the frame's rectangle gave 0 widget-coloured pixels against 107,776.
cost A payment form, a status board or a support widget is absent for every visitor while every capture and every DOM assertion says it is there. Because the parent page is intact, monitoring built on the parent stays green.
generalises Every boundary where the failure is declared by the far side and absorbed silently by the near one: blocked embeds, CORS-rejected fetches, refused redirects, sandboxed scripts.
mitigation Treat an embed as a dependency with its own health check: assert on the frame's load event or its resource entry, not on the presence of the element.
source https://www.rfc-editor.org/rfc/rfc7034
provenance documented
NS-073 An ancestor's overflow reassigns what a sticky element sticks to
reads as The header is declared `position: sticky; top: 0`, the capture shows it at the top of the page, and getComputedStyle reports `sticky`. Conclusion drawn: it sticks.
actually CSS Position 3 defines sticky as identical to relative except that its offsets are automatically adjusted in reference to the nearest ancestor scroll container's scrollport. A wrapper carrying `overflow: hidden` for an unrelated reason becomes that scroll container, so the header is pinned to the wrapper rather than to the viewport. The wrapper scrolls away with the page and takes the header with it.
blind because Scroll offset zero is the one position at which a working and a broken sticky element are in the same place, and it is the position every capture is taken at. The computed value does not separate them either: it reads `sticky` in both cases, because the declaration won the cascade and the defect lives in an ancestor.
CHECK Scroll and re-measure, rather than reading the declaration or the resolved value: `[0, 800, 2000].map(y => { scrollTo(0, y); return el.getBoundingClientRect().top; })`. Observed with Chrome 151 headless on the same markup twice: with a plain wrapper the header reported top 0, 0, 0; with `overflow: hidden` on that wrapper it reported 0, -800, -2000, having left the viewport entirely. `getComputedStyle(el).position` returned 'sticky' in both runs.
cost The navigation is unreachable on every long page, and the property is re-declared, re-prefixed and re-tested on the element while the ancestor that revoked it is never examined.
generalises Any property whose effect is decided by an ancestor rather than by the element declaring it: sticky and fixed positioning, stacking contexts, percentage heights, transforms creating containing blocks.
mitigation Walk the ancestors and check for a scroll container: any ancestor whose computed overflow is not `visible` in the sticky axis is the one the element is pinned to.
source https://www.w3.org/TR/css-position-3/#sticky-pos
provenance documented
NS-074 A full-page capture paints a fixed element once, at the offset it was captured from
reads as The full-page capture is 2400 pixels tall and the consent banner appears as a stripe near the top, well clear of the call to action further down. Conclusion drawn: the banner obstructs nothing.
actually A capture beyond the viewport renders the document once; the DevTools Protocol describes captureBeyondViewport as no more than capturing the screenshot beyond the viewport. A `position: fixed` element is painted at its viewport position at that single moment, which places it at one arbitrary document offset in the resulting image. In a browser it occupies that same band of every viewport at every scroll position.
blind because A full-page image is a picture in document space; a fixed element lives in viewport space. Flattening one onto the other destroys exactly the property that made the element worth checking, and leaves an image in which the overlay covers a small fraction of a very tall page.
CHECK Ask each fixed element what share of the viewport it owns: `[...document.querySelectorAll('*')].filter(e => getComputedStyle(e).position === 'fixed').map(e => e.className + ': ' + Math.round(e.getBoundingClientRect().height) + 'px = ' + Math.round(100 * e.getBoundingClientRect().height / innerHeight) + '% of every viewport')`. Observed with Chrome 151 headless: `["banner: 180px = 39% of every viewport"]`. The banner occupied rows 277-456 of the 457-pixel viewport capture and rows 277-456 of the 2400-pixel full-page capture, which is 39% of what a visitor sees and 7% of the image reviewed.
cost A banner, chat launcher or toolbar that permanently covers the bottom third of every screen is signed off against an image in which it covers a stripe of empty page, and the obstruction is only discovered from conversion data.
generalises Every rendering that resolves a relative frame into an absolute one: fixed positioning in full-page captures, relative timestamps baked into a report, `~` expanded at write time rather than read time.
mitigation Review fixed elements from viewport captures at several scroll offsets, and use `document.elementFromPoint` at the coordinates of anything that must remain clickable.
source https://chromedevtools.github.io/devtools-protocol/tot/Page/#method-captureScreenshot
provenance documented
======================================================================
INSTRUMENT: exit-code — A command's return status
======================================================================
Used when: The command exited zero, so you moved on.
Captures: Whether the process believed it completed the operation it chose to attempt.
Cannot see:
- Semantics. Zero means 'no error', never 'the thing you wanted is now true'.
- No-ops. A command whose behaviour depends on current state can succeed by doing nothing.
- Partial completion, where early steps left convincing artifacts before a later step aborted.
- Which target was acted on, when the argument resolved differently than intended.
- Commands that never return at all, which produce no exit code to inspect.
NS-005 enable --now does not restart an already-running unit
reads as The command exits zero and the service is active. Conclusion drawn: the new code is live.
actually systemctl enable --now starts a stopped unit. On a running one it is a no-op. The old process, with the old ExecStart, survives.
blind because Exit code zero and `active (running)` are true statements about the wrong process.
CHECK Compare the unit's ExecStart on disk against the live process: systemctl show -p ExecStart NAME and ps -p $MAINPID -o args=
cost A deploy is reported as complete twice while the previous binary keeps serving.
generalises Any idempotent-looking command whose semantics differ by current state.
provenance observed
NS-008 set -e aborts a script at a validation step that concerns something else
reads as The install script ran and the config file is in place. Conclusion drawn: the change is active.
actually A validation step covering the whole configuration failed on an unrelated block that needed an environment variable the script did not load. Under set -e the script exited before the reload.
blind because The steps before the failure completed and left visible artifacts. Partial success looks like success when only the artifacts are inspected.
CHECK Ask the running service what it loaded, not the filesystem what it holds. For Caddy: the admin API's live config.
cost The config is correct on disk and absent from the process, an inconsistency that survives inspection of either side alone.
generalises Any pipeline where a global check gates a local change.
provenance observed
NS-013 A teardown script destroys the environment it is executing inside
reads as Several long-running sessions vanish at once with no error output. Conclusion drawn: the tool crashed, or the machine failed.
actually A rebuild script ran `kill-session` against the multiplexer session it was itself running in. It killed its own parent, taking four unrelated sessions with it. There is no crash and no error because the script did exactly what it was told.
blind because A process that is killed cannot report that it was killed, and cannot report why. The absence of an error reads as an unexplained crash rather than a successful destructive command.
CHECK Before any teardown, compare the target against the environment you occupy: for tmux, test whether $TMUX is set and whether its session name equals the target. Refuse if they match.
cost Work in progress across every session in the environment, lost with no diagnostic trail.
generalises Any tool that can destroy a container, session, service, or host that it might itself be running inside.
provenance observed
NS-014 A privilege prompt with nowhere to appear hangs instead of failing
reads as A deploy step produces no output and does not return. Conclusion drawn: the operation is slow, or the network is stalling.
actually The command needed a password. There is no terminal to prompt on, so it waits indefinitely. No error, no exit code, no timeout.
blind because Exit codes only exist for processes that exit. An instrument that reads return status has nothing at all to read, and silence resembles work in progress.
CHECK Ask whether credentials are needed before running the real command: sudo -n true returns non-zero immediately when a password would be required.
cost An agent waits on a command that will never return, and a task that needed a human is reported as in progress.
generalises Every interactive prompt reached from a non-interactive context: credentials, confirmations, pagers, editors.
mitigation Wrap anything that might prompt in a timeout, so a hang converts into a failure you can observe.
provenance observed
NS-015 A pipeline returns the status of its last command, not its failing one
reads as `npm test | tee build.log` exits zero and the log file is written. Conclusion drawn: the tests passed.
actually A shell reports the exit status of the last command in a pipeline. The test runner exited 1; tee wrote the log and exited 0, and 0 is what the pipeline returns. `set -e` does not intervene, because the pipeline as a whole succeeded.
blind because One number is produced for a chain of processes. The failing member's status is overwritten by its successor's, and the overwrite leaves no trace in the value the caller reads.
CHECK Read the whole vector rather than the summary: `false | true; echo "${PIPESTATUS[@]}"` prints `1 0` where `$?` prints `0`. Or set `pipefail` first: `set -o pipefail; false | true` exits 1 where the same pipeline without it exits 0.
cost Every failure inside a command piped into tee, grep, jq, head or a formatter is recorded as success. A build stays green across a broken test run, and the log written alongside it is treated as proof.
generalises Any composition that collapses several results into one and keeps the last rather than the worst.
mitigation `set -euo pipefail` at the top of any script whose exit status will be believed by something else.
source https://www.gnu.org/software/bash/manual/bash.html#Pipelines
provenance documented
NS-016 curl exits zero after successfully downloading an error page
reads as `curl -s -o data.json URL` exits 0 and data.json exists with content in it. Conclusion drawn: the fetch succeeded.
actually The server answered 404 or 500. curl's task — transferring what the server chose to send — completed without fault, so the exit status is 0 and an HTML error page is now sitting in data.json under the name of the expected document.
blind because The exit code describes the transfer, not the response. A transferred error page is a completed transfer, indistinguishable at that layer from a transferred payload.
CHECK Ask for the status separately, or make curl care about it: `curl -s -o data.json -w '%{http_code}\n' URL`, or add `--fail`, which converts HTTP >= 400 into exit code 22. Observed on a 404: plain curl exits 0, `--fail` exits 22.
cost A downstream step parses an HTML error page as the config, dataset or credential file it expected. The failure surfaces at the parser, far from the request that caused it.
generalises Every client whose success criterion is that the protocol completed, rather than that the answer was the one asked for.
source https://curl.se/docs/manpage.html#-f
provenance documented
NS-017 A test that does not match the discovery pattern is neither run nor reported
reads as pytest exits 0 with a green summary after a new test is added. Conclusion drawn: the new test passes.
actually Collection matches `test_*.py` or `*_test.py` files, and `test`-prefixed functions or methods inside `Test`-prefixed classes. A file named `tests_auth.py`, or a function named `check_expiry`, is never collected. The green result belongs entirely to the other tests. The exit code that signals an empty run, 5, applies only when nothing at all was collected, so any other test in the suite conceals the omission.
blind because An uncollected test produces no pass line and no fail line. The summary counts what ran; it has no term for what was skipped by never being seen.
CHECK `pytest --collect-only -q | grep expiry` — prints the node id if the test was collected, prints nothing if it was not. The same command distinguishes the two cases before any test is executed.
cost The behaviour the test was written to protect is unprotected, and the suite's green status is subsequently cited as evidence that it is protected.
generalises Any convention-driven runner where registration is implicit and non-registration is silent: test discovery, plugin loaders, autoloaded fixtures, route decorators.
source https://docs.pytest.org/en/stable/explanation/goodpractices.html#conventions-for-python-test-discovery
provenance documented
NS-018 A bare mock answers to method names the real object no longer has
reads as The suite is green after a collaborator's method is renamed. Conclusion drawn: nothing depended on the old name.
actually `Mock()` manufactures an attribute on first access and returns another Mock, which is callable and truthy. Code calling `client.charge_card(...)` against the double passes although the real class now exposes only `charge`. The test exercises an interface that no longer exists, and will keep passing however far the real object drifts.
blind because The assertion is satisfied by the double's auto-created child. Green is a true statement about the mock, and the exit code cannot say which object the statement was about.
CHECK Derive the double from the real class: `create_autospec(Client)` or `Mock(spec=Client)` raises AttributeError on exactly the call a bare `Mock()` accepted. Observed on 3.12: `Mock().exsits()` returns a truthy Mock; `create_autospec(Real).exsits()` raises AttributeError.
cost A rename is shipped with a fully green suite whose coverage of the renamed path is zero. The regression appears in production, in code the tests appeared to cover.
generalises Every test double whose surface is invented rather than derived from the thing it replaces.
source https://docs.python.org/3/library/unittest.mock.html#autospeccing
provenance documented
NS-019 Outside strict mode MySQL stores an adjusted value and calls the statement successful
reads as The INSERT returns `Query OK, 1 row affected` and the client exits 0. Conclusion drawn: the row was stored as supplied.
actually With strict mode absent from sql_mode, MySQL 'inserts adjusted values for invalid or missing values and produces warnings'. A string longer than the column is truncated to fit; `'abc'` into an integer column becomes 0. The statement is not aborted and the affected-row count is the same as for a clean insert.
blind because Warnings are a separate channel that must be asked for. Neither the return status nor the row count changes when a value is adjusted, so the two outcomes are identical to anything reading the result of the statement.
CHECK `SHOW WARNINGS` (or `SHOW COUNT(*) WARNINGS`) immediately after the statement, in the same session: it returns rows such as `Data truncated for column ...` only when a value was adjusted, and nothing when it was not.
cost Truncated identifiers and coerced numbers are indistinguishable from real data once written, and the originals are gone. Corruption is discovered by a later join that finds nothing.
generalises Any writer that repairs input rather than rejecting it: lenient parsers, schema-on-read stores, spreadsheet imports.
mitigation Assert the mode rather than assume it: `SELECT @@SESSION.sql_mode` should contain STRICT_TRANS_TABLES before any load is trusted.
source https://dev.mysql.com/doc/refman/8.4/en/sql-mode.html#sql-mode-strict
provenance documented
NS-053 wait without arguments returns zero however its children exited
reads as A script starts several jobs with `&`, calls `wait`, and exits zero. Conclusion drawn: every job succeeded.
actually The bash manual is explicit: 'If id is not given, wait waits for all running background jobs and the last-executed process substitution, if its process id is the same as $!, and the return status is zero.' The children's statuses are reaped and discarded. Any number of them may have failed.
blind because An exit code reports what the last command chose to return, and bare wait returns zero by specification. The failing work happened in processes whose status was never requested.
CHECK Wait on each recorded PID and keep the statuses: `rc=0; for p in "${pids[@]}"; do wait "$p" || rc=$?; done; exit $rc`. Observed on bash 5.2.21 with one child exiting 3 and another exiting 7: bare `wait` returned 0, `wait $pid` on the second returned 7, and `wait -n` returned the status of the first job to finish.
cost Parallelism converts a failing step into a silent one. A fan-out of uploads, migrations or builds reports success while an arbitrary subset of it did not happen.
generalises Every aggregator that reduces many statuses to one: parallel test runners, job schedulers, batch APIs, fan-out without per-item bookkeeping.
source https://www.gnu.org/software/bash/manual/bash.html#Job-Control-Builtins
provenance documented
NS-054 Output redirection empties the file before the command reads it
reads as `sort data.txt > data.txt` exits zero and data.txt is still there. Conclusion drawn: the file was sorted in place.
actually The shell performs redirections before running the command, and for output redirection 'if the file does not exist it is created; if it does exist it is truncated to zero size'. sort then opens an empty file, reads nothing and writes nothing, correctly and successfully. The original contents are gone.
blind because The exit code belongs to a command that did exactly what was asked of it with the input it was given. The destruction happened in the shell, before the command started, and produced no status of its own.
CHECK Compare the line count before and after in the same command. Observed on bash 5.2.21: a three-line data.txt held zero lines after `sort data.txt > data.txt`, with sort exiting 0; `grep -v DEBUG conf.txt > conf.txt` left conf.txt at zero bytes, with grep exiting 1 because it had nothing to match.
cost The file that was supposed to be filtered is now empty, and emptiness is valid input to whatever reads it next: configuration becomes all-defaults, a dataset becomes zero records, and neither state raises an error.
generalises Every operation that opens its destination before reading its source: in-place archive rewrites, dumps piped over their own file, copies where source and destination alias.
mitigation Write to a new name and rename over the original, or use a tool with an explicit in-place mode such as `sed -i` or `sponge`.
source https://www.gnu.org/software/bash/manual/bash.html#Redirections
provenance documented
NS-055 find exits zero regardless of what the command it ran returned
reads as `find . -name '*.json' -exec validate {} \;` exits zero. Conclusion drawn: every file passed validation.
actually find's exit status describes find's own traversal. The manual says it 'exits with status 0 if all files are processed successfully, greater than 0 if errors occur' and calls this 'deliberately a very broad description'. The status of each -exec child is not part of it. All of them may have failed.
blind because One process walked the tree and a different process did the work. The status available to the caller belongs to the one that only walked.
CHECK Dispatch through a tool whose status covers the children: `find . -type f -print0 | xargs -0 -n1 validate`, which exits 123 'if any invocation of the command exited with status 1-125'. Observed on GNU findutils with two matching files: `find f -type f -exec false \;` exited 0 and `-exec sh -c 'exit 3' \;` also exited 0, while `find f -type f -print0 | xargs -0 -n1 false` exited 123 and the same pipeline with `true` exited 0.
cost A validation, conversion or upload sweep reports success across an entire tree while every item in it failed, and a sweep is precisely the step nobody re-checks item by item.
generalises Every dispatcher whose status covers dispatch rather than outcome: cron wrappers, CI steps that shell out, message producers acknowledged on enqueue.
source https://man7.org/linux/man-pages/man1/find.1.html
provenance documented
NS-056 A source path without a trailing slash adds a directory level at the destination
reads as `rsync -a build /var/www/site/` exits zero and the files are present under /var/www/site. Conclusion drawn: the build was deployed.
actually rsync's manual: 'A trailing slash on the source changes this behavior to avoid creating an additional directory level at the destination.' Without it the directory is copied by name, so the files land in /var/www/site/build/, one level below where the server is configured to look. The previous build continues to be served.
blind because The transfer succeeded and every file was copied correctly to a real path. The exit code describes the copy, not the destination's relationship to whatever reads it.
CHECK List the destination rather than trusting the status: `find /var/www/site -maxdepth 2 -name index.html`. Observed on rsync 3.2.7: `rsync -a rs/src rs/dest/` exited 0 and produced rs/dest/src/index.html, while `rsync -a rs/src/ rs/dest/` exited 0 and produced rs/dest/index.html.
cost The deploy reports success and the site does not change. Re-running it reproduces the same success, so the natural response to the symptom confirms the wrong hypothesis.
generalises Every copy whose destination semantics depend on a trailing character or on whether the target already exists: cp, scp, docker COPY, object-storage prefixes.
source https://man7.org/linux/man-pages/man1/rsync.1.html
provenance documented
NS-080 cd with an empty or unset argument succeeds without going anywhere
reads as `cd "$BUILD_DIR" && rm -rf ./*` exits zero and the script proceeds. Conclusion drawn: the build directory was entered and cleared.
actually bash(1): change the current directory to dir; if dir is not supplied, the value of the HOME shell variable is the default. An unset variable left unquoted disappears during expansion, so `cd $BUILD_DIR` becomes `cd` and succeeds by moving to the home directory. Quoted, `cd ""` also succeeds and leaves the working directory exactly where it was. Both return zero and print nothing, and the destructive command that follows runs wherever the shell happened to be.
blind because cd's status reports whether a directory change was performed, never which directory. Arriving where you intended and arriving in the home directory are the same value.
CHECK Confirm the destination rather than the status, or refuse an empty value outright: `: "${BUILD_DIR:?BUILD_DIR is empty}"; cd "$BUILD_DIR" && [ "$PWD" = "$BUILD_DIR" ]`. Observed on bash 5.2.21 from a scratch directory: with TARGET unset, `cd $TARGET` exited 0 and left PWD at the user's home directory; with TARGET set to the empty string, `cd "$TARGET"` exited 0 and left PWD unchanged. `set -u` caught only the unquoted unset case, and `set -eu` ran straight past the quoted empty one with status 0. With CDPATH=/usr, `cd bin` from /tmp exited 0 in /usr/bin.
cost The rm, rsync or build step that follows operates on the wrong tree with full confidence, and in the home-directory case on a tree that contains everything.
generalises Every command that treats a missing argument as a request for its default rather than as an error: cd, `git checkout`, `kubectl` without a namespace, `docker build` with an empty context path.
mitigation `${VAR:?}` fails on empty as well as unset, which `set -u` does not; assert on $PWD after any cd whose argument came from a variable.
source https://man7.org/linux/man-pages/man1/bash.1.html
provenance documented
NS-081 A declaration builtin consumes the exit status of the substitution it assigns
reads as `local token=$(fetch_token)` is followed by a status check, and the check passes. Conclusion drawn: fetch_token succeeded and token holds a token.
actually bash(1) explains the ordinary case: if no command name results and one of the expansions contained a command substitution, the exit status of the command is the exit status of the last command substitution performed. Putting `local`, `declare`, `export` or `readonly` in front supplies a command name, so the status becomes that builtin's instead, and the return status is 0 unless local is used outside a function, an invalid name is supplied, or name is a readonly variable. The substitution's failure is discarded and the variable holds an empty string.
blind because One line performs two operations and reports on the outer one. The status is a true statement about whether a variable was declared, offered where a statement about whether a value was obtained is expected.
CHECK Separate the declaration from the assignment and compare the two forms: `local token; token=$(fetch_token)`. Observed on bash 5.2.21: `f(){ local out; out=$(false); echo $?; }` printed 1, `g(){ local out=$(false); echo $?; }` printed 0, and `h(){ export OUT=$(false); echo $?; }` printed 0. Under `set -e` the split form aborted the shell and the combined form ran on to completion returning 0.
cost An empty credential, empty version string or empty path is carried forward and fails somewhere far from its origin, usually as an authentication error or a path that resolves to the filesystem root.
generalises Every wrapper that reports on itself rather than on what it invoked: shell builtins in front of assignments, test harnesses swallowing setup failures, entrypoints exiting on the shell's status rather than the program's.
mitigation Declare on one line and assign on the next wherever the command's outcome matters; `set -e` only helps once the two are separated.
source https://man7.org/linux/man-pages/man1/bash.1.html
provenance documented
NS-082 An unmatched pattern is passed through as a literal filename
reads as `for f in releases/*.tar.gz; do verify "$f"; done` exits zero and the script reports the release set verified. Conclusion drawn: every archive was checked.
actually bash(1): if no matching filenames are found, and the shell option nullglob is not enabled, the word is left unchanged. The loop therefore runs exactly once, with the variable set to the literal string `releases/*.tar.gz`, which names nothing. Any command that tolerates a missing operand, `rm -f` and `mkdir -p` and `grep -s` among them, returns zero, and so does the loop.
blind because Zero iterations and one iteration over a path that does not exist yield the same exit status. The difference lives in a count that nothing reports, and a wrong directory, a typo in an extension and an empty build all produce it identically.
CHECK Count what the pattern matched instead of what the loop returned: `shopt -s nullglob; files=(releases/*.tar.gz); echo "${#files[@]}"`, and fail on zero. Observed on bash 5.2.21 in an empty directory: `for fn in *.log; do echo "[$fn]"; done` printed `[*.log]` and exited 0; `rm -f *.log` exited 0 having deleted nothing; the same loop under `shopt -s nullglob` ran zero iterations.
cost A cleanup, upload or signing step reports success over an empty set, so the artefacts it was meant to handle survive untouched, and the next stage consumes the previous release without noticing it is stale.
generalises Every operation over a collection that is silently empty: globs, empty pipelines into xargs, queries returning no rows, iterations over an unset list.
mitigation Enable `nullglob` and assert on the array length, or `failglob` where an empty match is always a bug.
source https://man7.org/linux/man-pages/man1/bash.1.html
provenance documented
======================================================================
INSTRUMENT: http-response — A status code
======================================================================
Used when: You requested the URL and got 200.
Captures: That something on the path was willing to answer, and considered the answer complete.
Cannot see:
- Which layer answered: origin, CDN, proxy, or your own local cache.
- Whether the body is correct, or even the right document.
- Staleness. A cached copy returns 200 forever.
- Redirects, so the URL that answered may not be the URL you asked for.
NS-007 Heuristic HTML caching makes a completed deploy invisible to its author
reads as The old page is on screen after deploy. Conclusion drawn: the deploy failed.
actually Origin serves the new bytes. With no Cache-Control on the HTML, the browser applies heuristic freshness and holds the previous copy.
blind because The author is the least reliable observer of their own deploy, having visited the URL most often and therefore holding the strongest cache.
CHECK Hash the served bytes from a client that has never requested it: curl -s URL | md5sum, compared to the file on disk.
cost The deploy is repeated, rolled back, or rebuilt to fix a problem that exists only in one cache.
generalises Any layer that answers on the origin's behalf: CDN, proxy, DNS resolver, memoised client.
provenance observed
NS-020 S3 sends 200 OK before it knows whether the upload completed
reads as CompleteMultipartUpload returns HTTP 200. Conclusion drawn: the object is assembled and present.
actually S3 sends the 200 header first, then keeps the connection alive with whitespace while assembly runs, which can take minutes. A failure after that point is delivered as an `` document in the body of the response whose status line already said 200. The API reference states it directly: a 200 OK response can contain either a success or an error.
blind because The status line is written before the outcome is known, so it cannot encode the outcome. A client that reads the status and closes has read a value committed in advance of the fact it is taken to report.
CHECK Parse the body even on 200 and look for an `` root element; or confirm independently with HeadObject and compare ContentLength and ETag against what was uploaded. Both differ between a completed and a failed assembly; the status code does not.
cost An upload pipeline records success for an object that does not exist. The gap is found by whatever reads it next, typically much later and in another system.
generalises Any protocol that must acknowledge before it can know: streamed responses, long-polling, 202-style accepted work, write-behind caches.
source https://docs.aws.amazon.com/AmazonS3/latest/API/API_CompleteMultipartUpload.html
provenance documented
NS-021 A batch write returns 200 while handing back the items it did not write
reads as BatchWriteItem returns HTTP 200 and the SDK raises no exception. Conclusion drawn: all 25 items were written.
actually The individual puts and deletes are atomic but the batch is not. Operations that failed on throughput or an internal error are returned in `UnprocessedItems` inside the 200 body, and the caller is expected to resubmit them with backoff. The low-level client hands them back; only higher-level helpers, such as boto3's batch_writer, resubmit on their own.
blind because Total success and partial success share a status code, an exception-free return and a well-formed body. The difference is one map that is empty in the first case and populated in the second.
CHECK Assert the map is empty rather than assuming it: `sum(len(v) for v in resp.get('UnprocessedItems', {}).values()) == 0`. It is 0 on a full write and non-zero whenever items were dropped.
cost Rows go missing from a bulk load in proportion to how throttled the table was, with no error recorded anywhere, and the load is reported complete.
generalises Every bulk endpoint that reports transport success while carrying per-item failure in its payload.
source https://docs.aws.amazon.com/amazondynamodb/latest/APIReference/API_BatchWriteItem.html
provenance documented
NS-022 A single-page app's catch-all rewrite answers 200 for URLs that do not exist
reads as `curl -o /dev/null -w '%{http_code}' https://site/docs/pricing` returns 200. Conclusion drawn: the page exists and the link is good.
actually The host rewrites every unmatched path to index.html so the client-side router can handle it. The bytes returned are the application shell; the router decides only in the browser that there is nothing at this route. Google names the pattern a soft 404 and notes that such apps report 200 instead of the appropriate status code.
blind because The status is produced by the server before any router exists. Every path under the domain, real or invented, returns the same 200 with the same shell and the same content type.
CHECK Compare against a path that certainly does not exist: `curl -s $BASE/zzz-not-a-real-path | md5sum` and `curl -s $URL | md5sum`. Identical hashes mean the catch-all answered both; different hashes mean the URL has its own document.
cost Link checks, sitemap validation and 'the page is live' claims all pass against URLs with nothing behind them.
generalises Any fallback that answers on behalf of everything unmatched: wildcard DNS, default vhosts, permissive proxy routes.
source https://developers.google.com/search/docs/crawling-indexing/javascript/fix-search-javascript
provenance documented
NS-042 A truncated response arrives with its 200 already delivered
reads as `curl -s -o data.json -w '%{http_code}' URL` prints 200 and data.json holds plausible content. Conclusion drawn: the document was retrieved.
actually The status line is sent before the body. A connection that dies mid-body leaves a response RFC 9112 calls incomplete: 'a message body that uses the chunked transfer coding is incomplete if the zero-sized chunk that terminates the encoding has not been received', and a client that receives one 'MUST record the message as incomplete'. curl does record it, as exit code 18, but the status code stays 200 and the partial bytes stay on disk.
blind because %{http_code} is a property of the header, which was true when it was sent. Nothing checks length, because chunked encoding exists precisely for bodies whose length is not known when the header is written.
CHECK Read curl's exit status, not only the code it reports: `curl -s -o data.json -w '%{http_code}\n' URL; echo $?`. Observed against a local server that sends one chunk and closes the connection: `200` printed, exit status 18, and a 26-byte truncated prefix in data.json. `--fail` does not catch it, and `-s` suppresses the 'transfer closed with outstanding read data remaining' message that would otherwise appear.
cost A truncated CSV, NDJSON or log export is syntactically valid and merely shorter, so it loads without complaint and the missing records are indistinguishable from records that never existed.
generalises Every protocol that commits to a status before the payload is complete, and every consumer that validates syntax instead of completeness.
source https://www.rfc-editor.org/rfc/rfc9112#section-8
provenance documented
NS-043 curl -I asks a different question from the one users ask
reads as `curl -I https://site/report` returns 200 with a plausible Content-Length. Conclusion drawn: the page is up.
actually -I sends HEAD. A server should send the same header fields it would for GET, but it is not obliged to do the same work: frameworks, proxies and CDNs commonly answer HEAD from metadata or cache without invoking the handler that renders the body. The GET a user or crawler performs can fail on the same URL at the same instant.
blind because Status codes are per-method. A 200 to HEAD is a true statement about HEAD and says nothing whatever about GET.
CHECK Request the body and discard it: `curl -s -o /dev/null -w '%{http_code}\n' URL`. Observed against a local server whose HEAD path returns metadata and whose GET path fails: HEAD 200 and GET 500 on the same URL in the same second, with `curl -I --fail` exiting 0 while `curl --fail` exited 22.
cost Uptime monitors, link checkers and post-deploy smoke tests built on -I stay green throughout an outage every real request is experiencing.
generalises Every cheap probe standing in for the operation being assured: a TCP connect for a request, a ping for a service, a dry run for a run.
source https://www.rfc-editor.org/rfc/rfc9110#name-head
provenance documented
NS-044 curl drops the Authorization header when a redirect crosses to another origin
reads as `curl -sL -u user:pass https://host/private` returns 200 with a body. Conclusion drawn: the credentials were accepted and the private resource is reachable.
actually 'Authorization:' and 'Cookie:' headers are explicitly not passed on in HTTP requests when following redirects to other origins, unless --location-trusted is used. The first hop redirected elsewhere, curl reissued the request without the credentials, and the 200 belongs to whatever that origin serves anonymously: a login page, a public shell, or an empty result set.
blind because Every hop succeeded. The status code, the effective URL and the body are all consistent with an authenticated request, and the difference is a header removed by the client, which therefore appears in none of the server logs being consulted.
CHECK Ask the final host what it received, or compare the effective URL against the origin the credentials were issued for: `curl -sL -w '%{url_effective}\n' ...`. Observed with two local servers, the second echoing the header it received: `curl -sL -u alice:secret http://127.0.0.1:8933/data` printed 'AUTH RECEIVED: None' with status 200, as did the same request carrying an explicit `-H 'Authorization: Bearer ...'`; `--location-trusted` printed the Basic credential. 127.0.0.1 and localhost count as different origins.
cost An access-control change is signed off against a response that was never authenticated. The converse is worse: the same mechanism makes a broken authentication path look like a working one.
generalises Every hop that rewrites a request: proxies stripping headers, SDK retries losing context, message buses dropping attributes.
source https://curl.se/docs/manpage.html#--location-trusted
provenance documented
NS-045 A GraphQL endpoint returns 200 for a response whose data never arrived
reads as The POST returns HTTP 200 with a JSON body of the expected shape. Conclusion drawn: the query succeeded and the body holds the data.
actually Where the operation is executed and no request error is raised, the server should respond with 200 — 'this is the case even if a GraphQL field error is raised' during execution. Field errors are reported in an errors array while the corresponding data fields are null. Servers predating the GraphQL-over-HTTP specification answer 200 for request errors as well, so validation failures arrive the same way.
blind because The status code describes the transport; the outcome of the operation lives in the payload. A body carrying only errors is a well-formed 200 with the correct content type.
CHECK Assert on the payload: `jq -e 'has("errors") | not' resp.json`, and treat null leaves as failures rather than as absent data. Observed against the public countries.trevorblades.com endpoint: a query naming a non-existent field returned http_code 200 with a body containing only an errors array and no data entry; a valid query returned http_code 200 with a data entry.
cost A pipeline stores the null as a real value. The failure surfaces later as missing data, far from the query, with a 200 in the access log at the point where it actually went wrong.
generalises Every API carrying per-operation outcome in the body: JSON-RPC, batch endpoints, SOAP faults, webhook receivers that acknowledge before processing.
source https://graphql.github.io/graphql-over-http/draft/
provenance documented
NS-057 A response saved without decompression is stored as its compressed bytes
reads as `curl -H 'Accept-Encoding: gzip' -o data.json URL` reports 200, exits zero, and data.json is the expected size. Conclusion drawn: the document was fetched.
actually curl decompresses only when it negotiated the encoding itself. --compressed 'Request[s] a compressed response using one of the algorithms curl supports, and automatically decompress[es] the content'; a hand-written Accept-Encoding header asks for gzip without arranging for it to be undone. The file on disk begins 1f 8b and is a gzip member, not JSON.
blind because Status, exit code and transferred byte count are identical to a successful plain fetch. %{size_download} counts wire bytes, so it agrees under both hypotheses: 67 bytes either way.
CHECK Ask what the file is rather than how big it is: `file -b data.json`. Observed on curl 8.5.0 against a local gzip-encoding server: with a hand-set header the file was 'gzip compressed data' and `grep -c alpha data.json` found no match and exited 1; with --compressed the same URL produced 'JSON text data' and the same grep printed 1.
cost A search over the artifact returns nothing and the absence is read as a fact about the content. Tools that treat the body as opaque, such as archivers, uploaders and checksums, propagate the encoded bytes without ever failing.
generalises Every transport-level transformation the receiver must undo: content-encoding, transfer-encoding, base64 envelopes, client-side decryption.
mitigation curl's manual warns that saved response headers are not modified, so a stored header still claims the content is compressed after curl has decompressed it. The header is not a reliable record either way.
source https://curl.se/docs/manpage.html#--compressed
provenance documented
NS-058 A 200 carrying an Age header was answered by a cache, not by the origin
reads as `curl -sI https://site/` returns 200 and a Date header. Conclusion drawn: the origin is serving this, now.
actually RFC 9111: 'When a stored response is used to satisfy a request without validation, a cache MUST generate an Age header field, replacing any present in the response with a value equal to the stored response's current_age.' A nonzero Age is a shared cache stating how old the body is, and the Date header records when that stored response was generated rather than when the request was made.
blind because A status code does not identify the responder. Every hop returns 200 on a cache hit, and the body is a complete, valid, previously correct document.
CHECK Read the caching headers alongside the status: `curl -sI URL | grep -iE '^(date|age|x-cache|cf-cache-status):'`. Observed at 20:07:28 UTC: https://vercel.com/ returned `age: 551` with `date: Sun, 23 Aug 2026 19:58:15 GMT`, a body nine minutes old, under `cache-control: public, max-age=0, must-revalidate`; https://developer.mozilla.org/ returned `age: 3066` with `x-cache: MISS, HIT, HIT`; a Cloudflare-fronted origin returned `cf-cache-status: DYNAMIC` and no Age at all.
cost A deploy is verified against content that predates it, and the verification is stable: repeating the request returns the same stored copy with a larger Age, which reads as consistency.
generalises Every layer that may answer on another's behalf: CDNs, reverse proxies, resolvers, package mirrors, SDK-level response caches.
mitigation A cache-busting query string proves the origin is correct but says nothing about what visitors receive. Both readings are needed, and they answer different questions.
source https://www.rfc-editor.org/rfc/rfc9111#section-5.1
provenance documented
NS-059 A request issued from the origin host does not travel the path visitors take
reads as `curl -s -o /dev/null -w '%{http_code}' https://example.com/` run on the server returns 200. Conclusion drawn: the site is reachable and correct for visitors.
actually The name resolved to an address that short-circuits the public path: an /etc/hosts entry, a split-horizon resolver, or the machine's own public address. The origin answered directly, and the CDN, WAF, redirect rules and edge certificate that every visitor traverses were not involved. A failure in any of them is unreachable by this request.
blind because A status code records that something answered. Which of several layers answered is encoded nowhere in it, and a healthy origin behind a broken edge returns the same 200 as a healthy edge.
CHECK Record who answered, not just what: `curl -s -o /dev/null -w 'code=%{http_code} remote=%{remote_ip}\n' URL`. Compare that address against the origin you deployed to. A proxied domain answers from the proxy's address whether or not the origin behind it is alive; an unproxied one answers from the origin itself. Only the second reading tells you the origin is serving.
cost Edge misconfiguration, an expired certificate, a wrong origin rule or a route blocked by a WAF is confirmed working by a check structurally incapable of reaching it.
generalises Any probe issued from inside the system it measures: internal health checks, same-network monitoring, tests that resolve names through a private zone.
source https://curl.se/docs/manpage.html#-w
provenance documented
NS-075 A Content-Length shorter than the body truncates the response with no error anywhere
reads as `curl -s -o data.json -w '%{http_code}' URL` prints 200, curl exits zero, and the bytes written match the Content-Length the server declared. Conclusion drawn: the document was retrieved intact.
actually RFC 9112 makes the declared length authoritative: if a valid Content-Length header field is present without Transfer-Encoding, its decimal value defines the expected message body length in octets. The client reads that many and stops, whatever else is on the connection. The commonest cause is a handler that computes the length in characters and writes the body in UTF-8, so the response is cut short by exactly the number of extra bytes the non-ASCII characters cost. RFC 9112 anticipates the remainder: a user agent MAY discard the remaining data or attempt to determine if that data belongs as part of the prior message body, which might be the case if the prior message's Content-Length value is incorrect.
blind because Every length-based check agrees with itself. The header says 67, the file on disk is 67 bytes, and `%{size_download}` is 67. Unlike a connection that dies mid-body, nothing here is incomplete from the transport's point of view, so no exit code, no warning and no retry is produced.
CHECK Validate the body on its own terms rather than on the sender's: `curl -s URL | python3 -c 'import sys, json; json.load(sys.stdin)'`. Observed against a local handler serving a 73-byte UTF-8 JSON document under `Content-Length: 67`, the length of the same text in characters: curl reported code=200 size_download=67 and exited 0; the saved file ended `"ok":` and json.load raised `JSONDecodeError: Expecting value: line 1 column 62`. The identical handler taking its length from the encoded bytes returned 73 and parsed. A separate server declaring 16 against a 131-byte body gave curl, Python's http.client and Node all the same silent 16-byte prefix with status 200.
cost A record set is short by a few entries, a token is cut mid-string, a page loses its closing markup. Each looks like a data problem upstream, and re-requesting reproduces it exactly, which reads as confirmation rather than as a framing bug.
generalises Every length or checksum supplied by the same party that produced the payload, and every count taken in one unit and spent in another.
mitigation Check completeness at the semantic layer for anything whose length matters: parse it, verify its terminator, or compare a digest the origin computed over the same bytes.
source https://www.rfc-editor.org/rfc/rfc9112#section-6.3
provenance documented
NS-076 Two field lines with one name collapse into one value, and clients disagree about which
reads as `resp.headers['X-Frame-Options'] == 'DENY'` and the assertion passes. Conclusion drawn: the response carries the header the policy requires.
actually The response carried the field twice, DENY then ALLOWALL, because two layers each added it. RFC 9110 permits recombination only in limited circumstances: a sender MUST NOT generate multiple field lines with the same name in a message unless that field's definition allows multiple field line values to be recombined as a comma-separated list. Senders do it anyway, and recipients then differ. Some return the first value, some the joined list, some keep both. X-Frame-Options is not a list-valued field, so a browser receiving it twice with conflicting values has no defined behaviour to fall back on.
blind because A header dictionary maps one name to one string. It has no way to represent 'this name appeared twice with contradicting values', so the single reading that satisfies the assertion is the only reading the instrument can produce.
CHECK Read the field lines rather than the parsed mapping: `curl -sD - -o /dev/null URL | grep -ci '^x-frame-options:'`, and treat any count above one as a failure. Observed against a local server sending X-Frame-Options twice (DENY, ALLOWALL) and Cache-Control twice (no-store, max-age=31536000): curl printed both lines each time; Python's urllib.request returned 'DENY' and 'no-store' from `headers[name]` while `headers.get_all` returned both values; `http.client.getheader` returned the joined 'no-store, max-age=31536000'; Node 20.20.2 returned the joined 'DENY, ALLOWALL'. One response, three clients, three different answers.
cost A security-header audit passes on a response no browser will honour, and a caching audit reads no-store while a shared cache reads the joined value and stores the response for a year.
generalises Every flattening of a multi-valued source into a scalar: repeated query parameters, environment variables set twice, merged configuration layers, duplicate rows collapsed by a join.
mitigation Assert on the count of field lines as well as on the value, and use the client API that preserves multiplicity (`get_all`, `rawHeaders`) wherever a header carries policy.
source https://www.rfc-editor.org/rfc/rfc9110#section-5.3
provenance documented
NS-077 Overriding the Host header leaves the TLS handshake pointing somewhere else
reads as `curl -H 'Host: app.example.com' https:///` returns 200 with a plausible page. Conclusion drawn: the origin serves app.example.com correctly.
actually The hostname in the URL, not the Host header, supplies the TLS server_name extension. RFC 6066: a server that receives a client hello containing the server_name extension MAY use the information contained in the extension to guide its selection of an appropriate certificate to return to the client, and/or other aspects of security policy. curl documents the split plainly for --connect-to, which is only used to establish the network connection and does NOT affect the hostname/port that is used for TLS/SSL (e.g. SNI, certificate verification). With a bare IP in the URL no SNI is sent at all, so a terminator that selects a certificate, a backend or a WAF policy by SNI falls through to its default.
blind because The 200 is real and the body may well be the right site, because the default virtual host is often the site under test. Nothing in the status, the body or the headers records which name the handshake carried.
CHECK Keep the hostname in the URL and move only the address: `curl -sk --resolve app.example.com:443: https://app.example.com/`. Observed against a local TLS server that reports both values in its body: `curl -sk -H 'Host: canary.example' https://127.0.0.1:8443/` returned `SNI=None HOST=canary.example`, while `curl -sk --resolve canary.example:8443:127.0.0.1 https://canary.example:8443/` returned `SNI=canary.example HOST=canary.example:8443`. `openssl s_client` split the same way: no -servername, no SNI.
cost A certificate, a routing rule or a security policy bound to the hostname is never exercised, so the probe passes for a name that would fail for every real client, and the failure appears only when DNS is cut over.
generalises Every request whose identity is asserted at more than one layer: SNI against Host, DNS name against certificate subject, connection string host against database catalogue, tenant header against signed token.
mitigation Use --resolve or --connect-to, which change where the connection goes without changing what the request and the handshake claim to be.
source https://www.rfc-editor.org/rfc/rfc6066#section-3
provenance documented
NS-078 The first status line in a response may belong to an interim response
reads as `curl -sD headers.txt URL` succeeded and headers.txt begins with a status line followed by a header block. Conclusion drawn: those are the response's status and its headers.
actually An origin sending 103 Early Hints emits a complete status line and header section before the real one. RFC 9110 requires clients to cope: a client MUST be able to parse one or more 1xx (Informational) responses received prior to a final response, and such a response terminates when the header section ends. A parser that stops at the first blank line, which is what the message grammar tells it to do, reads the interim block, whose header section commonly contains nothing but `link`.
blind because Both blocks are well-formed HTTP with the same shape. Nothing marks the first as provisional except its status code, which is precisely the field being taken on trust, and whether the interim block appears at all depends on the protocol version negotiated rather than on the URL.
CHECK Count the status lines before reading any of them: `curl -sD - -o /dev/null --http2 URL | grep -c '^HTTP/'`, and treat anything above one as two blocks to disentangle. Observed at 00:28 UTC on 2026-08-24 against https://www.cloudflare.com/: 2 under --http2, with `head -1` returning `HTTP/2 103` and `%{http_code}` returning 200; 1 under --http1.1, where the same origin sent only `HTTP/1.1 200 OK`.
cost A header audit reads the Early Hints block, finds one `link` field, and reports HSTS, CSP, Content-Type and Cache-Control as absent from a response that carries all four. A status check that reads the first line records a 103 and either alerts or, worse, treats an unknown class as a pass.
generalises Every stream where a provisional message precedes the real one: 1xx responses, redirect chains, retried requests, partial results emitted before a final aggregate.
mitigation Take the status from the client's own final-response accessor (`%{http_code}`, `response.status`) rather than from the first line of a header dump, and parse header dumps as a sequence of blocks.
source https://www.rfc-editor.org/rfc/rfc9110#section-15.2
provenance documented
NS-079 Field names arrive lowercased over HTTP/2, so a case-sensitive check passes vacuously
reads as The audit greps the response headers for `^Set-Cookie:` lines lacking the Secure attribute, finds none, and records a pass. Conclusion drawn: no insecure cookie is being set.
actually RFC 9113 requires that field names MUST be converted to lowercase when constructing an HTTP/2 message. Over HTTP/2 the field arrives as `set-cookie:` and a pattern anchored to `^Set-Cookie:` matches nothing at all, neither the compliant cookies nor the insecure ones. Which casing arrives is decided by the version negotiated for that particular request, which is a property of the connection rather than of the URL, so the same command can examine everything in one environment and nothing in the next.
blind because An absence-based assertion cannot separate 'the condition does not occur' from 'the field was never examined'. Both return zero matches, exit the same way, and read as a clean result.
CHECK Prove the pattern matches something before trusting that it matched nothing, and record the version alongside it: compare `curl -sD - -o /dev/null --http1.1 URL | grep -c '^Content-Type:'` against the same command with `--http2`. Observed at 00:28 UTC on 2026-08-24 against https://www.cloudflare.com/: 1 under --http1.1, 0 under --http2, and 1 under --http2 for `grep -c '^content-type:'`. `--http2` had also negotiated 1.1 without complaint against a local HTTP/1.1-only server, reporting `version=1.1 code=200` and exiting 0.
cost A control that has never once evaluated its subject reports a pass on every run, and the run history makes it look like a stable, long-standing green.
generalises Every check whose passing condition is an empty result set: greps for a forbidden pattern, queries returning no rows, log searches finding no errors, linters with a misconfigured include path.
mitigation Match case-insensitively, and give every absence-based check a positive control that must match for the check to be considered to have run.
source https://www.rfc-editor.org/rfc/rfc9113#section-8.2
provenance documented
======================================================================
INSTRUMENT: file-on-disk — The file's contents
======================================================================
Used when: You read the config, the stylesheet, or the source and confirmed it says the right thing.
Captures: Authored intent, at the moment you read it.
Cannot see:
- What a running process actually loaded, which may predate your edit.
- What overrode it later, in any last-writer-wins system.
- Whether the file is the one being served, as opposed to one with the same name elsewhere.
- Whether a plausible-looking artifact is the correct artifact.
NS-003 A later cascade rule silently revokes position: fixed
reads as The element is declared `position: fixed` in the stylesheet. Conclusion drawn: the browser or the viewport is at fault.
actually A rule added later for an unrelated purpose (`main, .nav, .colophon { position: relative }`) matched the same element and won on document order.
blind because Reading the intended declaration confirms the intent, never the outcome. The stylesheet says fixed; the element is relative.
CHECK getComputedStyle(el).position — the resolved value, never the authored one.
cost The property is re-declared, re-prefixed and re-tested while the overriding rule stays untouched.
generalises Any last-writer-wins system: CSS, environment variables, layered configuration, merged dictionaries.
provenance observed
NS-006 Scraping the largest asset returns a recommendation, not the subject
reads as A parser extracts an image from the page and it is a valid, plausible image. Conclusion drawn: extraction succeeded.
actually The page embeds related items alongside the subject. Ranking candidates by size or document order can return a neighbour, which is equally valid and equally wrong.
blind because Both results are real images from the correct domain. Nothing about the artifact reveals it is the wrong one.
CHECK Prefer the canonical marker the page declares about itself (og:image, canonical link, structured data) over any heuristic ranking of candidates.
cost Silent substitution. Detected only when two different inputs return the same output.
generalises Any extraction from a document that also describes things other than itself.
provenance observed
NS-023 sshd takes the first value for a keyword, so an appended directive loses to an include
reads as /etc/ssh/sshd_config ends with `PasswordAuthentication no`, and sshd reloaded without error. Conclusion drawn: password logins are disabled.
actually The man page states that unless noted otherwise, for each keyword the first obtained value will be used. On a stock Ubuntu image `Include /etc/ssh/sshd_config.d/*.conf` sits at line 12 of a 131-line file, so a drop-in such as 50-cloud-init.conf that sets the same keyword is read first and wins. The line appended at the bottom is parsed and discarded.
blind because The file says what was intended, and it is the file that was edited. Precedence is a property of the merge order across several files, and the include that pre-empts the edit sits above it, out of the region being read.
CHECK `sudo sshd -T | grep -i passwordauthentication` prints the effective merged value the daemon will use, which differs from the authored line whenever an earlier occurrence won. Run it with root privileges: as an unprivileged user it silently omits unreadable drop-ins.
cost A hardening change is recorded as applied while the setting it was meant to change is untouched, and the evidence for the claim is the file that lost.
generalises Every first-wins configuration system, which fails in exactly the opposite direction to the last-wins ones and therefore defeats the habit built on them.
source https://man.openbsd.org/sshd_config.5
provenance documented
NS-024 Bidirectional control characters make source read differently than it compiles
reads as A reviewer reads the diff and the early return is plainly inside a comment. Conclusion drawn: the change is inert.
actually Unicode bidirectional overrides (U+202A to U+202E, U+2066 to U+2069) reorder the display of tokens without changing their logical order. Compilers and interpreters adhere to the logical ordering of source code, not the visual order, so the code executed is not the code rendered. Catalogued as CVE-2021-42574, with a homoglyph variant as CVE-2021-42694.
blind because Reading a file means reading a rendering of it. The terminal, the editor and the diff viewer all apply the same bidi algorithm as the attack, so the instrument and the exploit agree with each other and disagree with the compiler.
CHECK Search for the characters instead of reading the text: `grep -rlP '[\x{202A}-\x{202E}\x{2066}-\x{2069}]' path/` names files containing them and prints nothing for files that do not. Verified against a planted sample and a clean file.
cost Code is reviewed and approved on the strength of behaviour no reviewer ever saw.
generalises Any check performed on a rendering of an artifact rather than on its bytes.
mitigation Compilers now detect this where asked: rustc's text_direction_codepoint_in_literal lint and gcc's -Wbidi-chars. Enable them rather than relying on reading.
source https://trojansource.codes/
provenance documented
NS-025 A JSON integer above 2^53 is silently rounded when parsed as a double
reads as The response contains `"id": 10765432100123456789`; the parsed object has an id of the right shape and it round-trips through the code. Conclusion drawn: the identifier was carried through intact.
actually JavaScript parses JSON numbers as IEEE 754 doubles. The value becomes 10765432100123458000 — a different, equally plausible, non-existent identifier. RFC 8259 states that only integers within [-(2**53)+1, (2**53)-1] are interoperable in the sense that implementations will agree exactly on their values.
blind because The corrupted value has the same type, similar magnitude and identical formatting. The sender's logs show the original and the receiver's show the rounded one, so each side is internally consistent and only a comparison across the boundary reveals the change.
CHECK `Number.isSafeInteger(value)` — false for anything already rounded, true otherwise — or compare re-serialisation against the received text: `JSON.stringify(JSON.parse(s)) === s`. Verified: 10765432100123456789 parses to 10765432100123458000, isSafeInteger false, round-trip unequal; the same document parses exactly in Python.
cost Reads and writes land on the wrong record or on none. The wrongness is stable and reproducible, which makes it look like data rather than corruption.
generalises Every boundary between systems with different numeric ranges: 64-bit ids into doubles, timestamps into 32-bit seconds, decimals into floats.
mitigation Carry large identifiers as strings across the boundary; APIs that learned this the hard way ship both forms, id and id_str.
source https://www.rfc-editor.org/rfc/rfc8259#section-6
provenance documented
NS-046 An unquoted YAML scalar becomes a boolean before anything reads it
reads as config.yml reads `country: NO` and `version: 1.10`, and it is the file the service loads. Conclusion drawn: those are the values the service has.
actually The YAML 1.1 boolean resolver matches y, yes, n, no, true, false, on and off in any capitalisation, and the loaders in widest use implement it. `NO` loads as False, `off` as False, `on` as True. Numeric resolution is equally implicit: `1.10` becomes the float 1.1 and `0xdeadbeef` becomes 3735928559.
blind because Reading the file confirms the characters, and the characters are right. The conversion happens inside the loader, and the loaded value is never displayed beside the text it came from.
CHECK Load it and print the types instead of reading it: `python3 -c "import yaml,sys;[print(repr(k),repr(v),type(v).__name__) for k,v in yaml.safe_load(open(sys.argv[1])).items()]" config.yml`. Observed on PyYAML 6.0.1: `NO` -> False, `off` -> False, `on` -> True, `1.10` -> 1.1 (float), `0xdeadbeef` -> 3735928559 (int), while `08` stayed the string '08' because it is not a valid octal literal — so neighbouring keys in one file resolve inconsistently.
cost A country code becomes a boolean and a version becomes a different version. Both are valid values of the wrong type, and both survive any review conducted by reading the file.
generalises Every format that infers type from spelling: spreadsheet imports turning identifiers into dates, shell word-splitting, environment variables parsed as numbers.
mitigation Quote every scalar whose type matters. Loaders following the YAML 1.2 core schema resolve fewer of these, which changes the set of surprises rather than removing it.
source https://yaml.org/type/bool.html
provenance documented
NS-047 A directory containing a .git is committed as a pointer rather than as files
reads as The files are in the working tree, `git add -A` and `git commit` both succeeded, and `git status` reports a clean tree. Conclusion drawn: the files are in the repository.
actually A directory with its own .git is recorded as a gitlink — one index entry of mode 160000 holding a commit id — not as the files beneath it. The commit id names an object in a repository nobody else can reach, and no .gitmodules entry is created. A clone contains the path as an empty directory.
blind because git status compares the working tree with the index, and the index is satisfied: the gitlink matches the nested repository's HEAD. Everything tracked is up to date, and the untracked files sit behind a boundary status does not cross. The warning appears once, at add time, on stderr, and does not affect the exit status.
CHECK Ask whether the specific file is tracked: `git ls-files --error-unmatch vendor/widget/index.js` — prints the path and exits 0 when it is, prints "did not match any file(s) known to git" and exits 1 when it is not. Observed here: `git ls-tree HEAD vendor/` returned `160000 commit bdf3631e... vendor/widget`, and a fresh clone of the repository contained README.md and nothing else.
cost The backup, the mirror or the deploy artefact is missing a subtree that every local check confirms is present. It is discovered on a clean clone, usually on another machine, usually once the original is gone.
generalises Every container that stores a reference where the reader assumes contents: symlinks inside archives, submodule pointers, lockfiles naming versions that no longer resolve.
source https://git-scm.com/docs/gitsubmodules
provenance documented
NS-048 The module that imports is the first file on the path with that name
reads as config.py was edited, saved, and re-read to confirm the new value; the program still uses the old one. Conclusion drawn: a caching problem, or the process was not restarted.
actually The interpreter searches sys.path in order, with the directory containing the input script placed at the front. Any file of that name in the script's directory, in the working directory, or earlier in the path is imported instead. The edited file is never read, and nothing is raised because the module that was found is a perfectly valid module.
blind because The file being read and the file being imported have the same name and a similar shape, and the import statement names neither directory. Nothing in the source distinguishes them.
CHECK Ask the imported module where it came from, invoked exactly as the program is invoked: `python3 -c 'import config; print(config.__file__)'`. Observed here with two config.py files present: a script in sub/ loaded sub/config.py and reported that path, while the copy that had been edited sat one directory up, untouched.
cost Edits accumulate in a file the program has never loaded, and the conclusion drawn concerns caching or process lifetime rather than identity, sending the work into restarts and clearing __pycache__.
generalises Every ordered resolution path where names are not unique: PATH, LD_LIBRARY_PATH, node_modules resolution, classpath, include directories.
source https://docs.python.org/3/tutorial/modules.html#the-module-search-path
provenance documented
NS-060 git add says nothing when a pathspec matches only ignored files
reads as `git add .` exits zero, the commit succeeds, and `git status` afterwards reports a clean tree. Conclusion drawn: everything in the working directory is committed.
actually git-add documents both branches: 'The git add command will not add ignored files by default. You can use the --force option to add ignored files. If you specify the exact filename of an ignored file, git add will fail with a list of ignored files. Otherwise it will silently ignore the file.' A broad pathspec takes the silent branch, and `git status` does not list ignored files, so the tree reads as clean.
blind because Both instruments agree, and both are answering a narrower question than the one asked. `git status` compares the index against the tracked working tree; a file that is neither tracked nor reportable is outside that comparison by construction.
CHECK Ask whether a specific path is excluded, and list what was excluded: `git check-ignore -v path` and `git status --short --ignored`. Observed on git 2.43.0 with a .gitignore containing dist/, *.local and config/*: `git add .` exited 0, `git status --short` listed only .gitignore and app.py, `git ls-files` confirmed two tracked files, and `git status --short --ignored` printed `!! config/`, `!! dist/` and `!! settings.local` for the three that were never staged.
cost A build output, a migration or a generated asset that a broad rule happens to match is absent from every clone and every deploy, and each step in the chain reports the tree as clean.
generalises Every filter applied before a report is produced: exclude rules in backups and syncs, packaging manifests, .dockerignore, test collection patterns.
source https://git-scm.com/docs/git-add
provenance documented
NS-061 Two readers of one CRLF file disagree about where each value ends
reads as cat .env prints `API_URL=https://api.example.com` and that is the file the service loads. Conclusion drawn: the service has that URL.
actually The file uses CRLF terminators. Python's default text mode translates them, so a Python reader sees a 23-character value; the shell does not translate, so `. ./.env` yields a 24-character value ending in a carriage return. The same bytes become different strings depending on who reads them.
blind because Reading the file with cat, an editor or a code review renders the carriage return as nothing at all. It has no glyph, occupies no column, and is removed by several of the tools used to inspect it.
CHECK Make the terminators visible, or measure the value in the reader that matters: `cat -A .env`. Observed on this host: cat -A printed `API_URL=https://api.example.com^M$`, `file` reported 'ASCII text, with CRLF line terminators', bash reported ${#API_URL} as 24 and the equality test against the intended URL failed, while Python text mode reported 23 and the same file opened in binary mode yielded 'https://api.example.com\r'.
cost A hostname, token or path acquires an invisible trailing byte. Comparisons fail, signatures do not verify, and a request goes to a name that differs from the one on screen, while every review of the file keeps confirming the value is right.
generalises Every character that renders as nothing: byte-order marks, zero-width spaces, non-breaking spaces pasted from documents, trailing whitespace in secrets.
mitigation Normalise on ingest with a gitattributes `text` rule, and compare lengths rather than appearances when a value refuses to match something it visibly equals.
source https://git-scm.com/docs/gitattributes
provenance documented
NS-062 Copying a symlink in archive mode produces a second link, not a backup
reads as `cp -a app.conf app.conf.bak` exits zero and a listing shows both names. Conclusion drawn: the original is preserved, so the edit is safe.
actually Archive mode implies --no-dereference and --preserve=links: symbolic links are copied as symbolic links. app.conf was a link, so app.conf.bak is a second link to the same target. There is one file. Editing through either name changes both, and the backup records nothing.
blind because Reading either path returns the intended contents, and a listing shows two entries with the expected names. Only the link marker and the inode number distinguish a backup from an alias.
CHECK Compare inodes after dereferencing: `stat -Lc '%i %n' app.conf app.conf.bak`. Observed on GNU coreutils: after `cp -a app.conf app.conf.bak` both names and the underlying real.conf reported inode 2142629, and overwriting app.conf with new content changed the contents visible through app.conf.bak at the same moment.
cost The rollback path does not exist, and its absence is discovered only when it is needed. A policy requiring a backup before editing is satisfied on paper by an operation that made none.
generalises Every duplication that may preserve a reference instead of the data: hard links, copy-on-write clones, container image layers, object-store copies that alias.
mitigation `cp -L` copies the target's contents. Checking the inode immediately afterwards costs one command and is the only cheap moment to find this.
source https://www.gnu.org/software/coreutils/manual/html_node/cp-invocation.html
provenance documented
NS-083 An in-place edit of a symlink replaces the link with a regular file
reads as `sed -i 's/old/new/' app.conf` exits zero and reading app.conf shows the new value. Conclusion drawn: the configuration was updated.
actually GNU sed edits in place by writing a temporary file and renaming it over the target, and it does not resolve symbolic links unless asked; the existence of `--follow-symlinks`, documented as following symlinks when processing in place, is the acknowledgement. The rename replaces the link itself. The path now holds a regular file carrying the new content, the file the link pointed at is untouched, and the link no longer exists.
blind because Reading the path returns the edited content, because the path genuinely holds it now. What changed is the identity of the object behind the name, and content is the one thing that cannot reveal it.
CHECK Look at the type and inode behind the name rather than at the bytes: `stat -c '%i %F %N' app.conf`. Observed on GNU sed 4.9 with app.conf a symlink to repo/app.conf: beforehand `2659981 symbolic link 'app.conf' -> 'repo/app.conf'`; after `sed -i`, `2659983 regular file 'app.conf'` holding `setting=new`, while repo/app.conf still held `setting=old` at its original inode 2659980. With `--follow-symlinks` the link survived and repo/app.conf received the edit.
cost The change is invisible to the repository the link came from, is not committed, and is silently reverted the next time the link farm is rebuilt or the host is reprovisioned. Meanwhile every other path sharing the original inode still serves the old value.
generalises Every write-by-rename: editors saving atomically, `sed -i`, `sort -o`, dotfile farms, hard links and bind mounts that expected to keep sharing an inode.
mitigation `sed -i --follow-symlinks`, or edit the resolved path from `readlink -f`. The same applies to any tool that writes by rename.
source https://man7.org/linux/man-pages/man1/sed.1.html
provenance documented
NS-084 A file that changed without changing size or timestamp is never transferred
reads as `rsync -a src/ dst/` exits zero, and dst/app.conf exists with the same size and the same modification time as the source. Conclusion drawn: the destination is a copy of the source.
actually rsync(1): rsync finds files that need to be transferred using a 'quick check' algorithm (by default) that looks for files that have changed in size or in last-modified time. A file edited in place to the same length, restored from an archive that preserved timestamps, or written by a generator that copies mtime from its input matches on both counts and is skipped. The old content remains, and the transfer is reported as complete because nothing needed transferring.
blind because The two numbers the tool decides on are the two numbers used to verify it. Size and mtime agree, which is a true statement about the metadata and a false one about the contents.
CHECK Compare contents rather than metadata: `rsync -ain --checksum src/ dst/` lists exactly what a content comparison would move. Observed on rsync 3.2.7 with src/app.conf holding VERSION=2 and dst/app.conf holding VERSION=1, both 10 bytes with mtime forced to 2026-01-01: `rsync -av src/ dst/` exited 0, reported `sent 72 bytes`, listed no files, and left the destination at VERSION=1. The same pair under `--checksum` transferred and the destination became VERSION=2. `cp -u` copied nothing for the same reason.
cost A deploy or backup succeeds and changes nothing, and repeating it reproduces the same success, so the obvious response to the symptom confirms the wrong hypothesis.
generalises Every change detector keyed on a proxy for content: mtime-based build systems, ETags derived from metadata, cache keys built from a version string that was not bumped.
mitigation Use `--checksum` for any sync where the content is authoritative and the timestamps are not, accepting the read cost; or ensure the writer touches mtime whenever it rewrites.
source https://man7.org/linux/man-pages/man1/rsync.1.html
provenance documented
NS-085 A repeated key is resolved silently, and the occurrence you read is not the one in force
reads as config.json declares `"debug": false` and a database host of prod.db.internal, and it is the file the service loads. Conclusion drawn: debug is off and the service talks to production.
actually The same names appear again further down the file, added by a later edit or a careless merge. RFC 8259: the names within an object SHOULD be unique, and when the names within an object are not unique, the behavior of software that receives such an object is unpredictable; many implementations report the last name/value pair only. The values in force are the ones at the bottom.
blind because Reading a configuration file means reading downwards and stopping at the first occurrence of the key in question. Nothing in the syntax marks a name as later overridden, and both occurrences are individually valid.
CHECK Load it through the parser the service uses and print what it produced: `python3 -c 'import json, sys; print(json.load(open(sys.argv[1])))' config.json`. Observed on Python 3.12.3, Node 20.20.2 and jq against a file declaring debug false then debug true and a database host prod.db.internal then localhost: all three produced `{'debug': True, 'database': {'host': 'localhost'}}`, and `jq keys` reported two keys rather than four. PyYAML 6.0.1 behaved the same way on the YAML equivalent. Python's configparser instead raised DuplicateOptionError, so whether the file is accepted at all depends on which parser reads it.
cost Debug output, a staging database or a permissive CORS origin is live in production, and the file that proves otherwise is the same file that enables it.
generalises Every last-writer-wins merge that leaves both writers visible: duplicate keys, repeated environment assignments, layered configuration, stylesheet rules of equal specificity.
mitigation Reject duplicates at load time; Python's `json.load` accepts an `object_pairs_hook` that can raise on a repeated name, and most YAML loaders can be configured to do the same.
source https://www.rfc-editor.org/rfc/rfc8259#section-4
provenance documented
======================================================================
INSTRUMENT: process-list — What is running
======================================================================
Used when: You checked ps, pgrep, or systemctl status.
Captures: That a process with a matching name exists and has a state.
Cannot see:
- Whether the running process corresponds to the code currently on disk.
- Your own command, which matches the pattern you are searching for.
- Whether the environment you are inspecting is the one you are running inside.
NS-012 A pattern search for a process matches the search itself
reads as pgrep -f chromium returns a match after cleanup. Conclusion drawn: an instance is still running.
actually The pattern appears in the command line of the pgrep invocation, so the search finds itself. Nothing is running.
blind because The instrument is a process, and it is inside the set it is measuring. The output format gives no indication which row is the observer.
CHECK Resolve the PID and compare it against your own: pgrep -f PATTERN | grep -v "^$$\$", or list full command lines with pgrep -af and read them.
cost Cleanup loops that never terminate, or a kill aimed at the shell performing the kill.
generalises Any measurement taken from inside the population being measured.
provenance observed
NS-026 systemd reports a Type=simple unit active before the service binary has been executed
reads as `systemctl start app` returns and `systemctl is-active app` says active. Conclusion drawn: the service is up and accepting connections.
actually For Type=simple the service manager considers the unit started immediately after the main service process has been forked off — after fork() and before the new process has called execve() to invoke the actual service binary. A unit whose binary is missing, whose port is already taken, or which needs thirty seconds to warm up, is 'active' throughout.
blind because The process list reports existence and state. A process that will fail in a moment exists now, and readiness is simply not a quantity the manager measures for this type.
CHECK Ask the socket rather than the manager: `ss -ltnp 'sport = :8000'` returns a listener only when one exists, and is empty while the unit is active but not yet serving. A single request to the port distinguishes the same two states.
cost Dependent units and deploy scripts proceed against a service that is not listening. The ordering guarantee that was assumed was never offered.
generalises Every start-up API that acknowledges the request rather than the readiness.
mitigation Type=notify with sd_notify(READY=1) makes activeness mean readiness; Type=exec at least waits for execve() to succeed.
source https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Type=
provenance documented
NS-027 A service crash-looping every few seconds reads as active between crashes
reads as `systemctl status app` shows active (running) with a PID. Conclusion drawn: the service is healthy.
actually With Restart=always the unit crashes, waits RestartSec, and starts again. Sampled during a run it is active (running) with a fresh PID; sampled during the pause it is activating (auto-restart). Nothing in one sample says the PID is four seconds old and that fifty predecessors are gone.
blind because The process list is a snapshot. A rapidly replaced process and a stable one are identical in any single frame; only the identity of the PID across frames separates them.
CHECK Read the restart counter and the start timestamp twice, thirty seconds apart: `systemctl show -p NRestarts -p ExecMainStartTimestamp --value app`. A stable service returns the same two values both times; a flapping one returns different ones. Both properties are exposed by systemd for every service unit.
cost A deploy is signed off on a service that drops every request arriving inside its restart window, until the start rate limit is reached and it stays down for good.
generalises Every supervised process where the supervisor's diligence in restarting is read as the process's success in running.
source https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Restart=
provenance documented
NS-036 A terminated process keeps its entry in the table until the parent reaps it
reads as `pgrep worker` still returns a PID after the shutdown request. Conclusion drawn: the worker is refusing to exit and shutdown is hanging.
actually The process has already exited. Its entry persists in state Z — what ps calls a 'defunct ("zombie") process, terminated but not reaped by its parent' — retaining only its exit status. It has no address space, executes nothing and cannot be killed; SIGKILL to a zombie does nothing. It disappears when the parent calls wait(), or when the parent itself exits and init reaps it.
blind because The process list reports existence. A zombie exists as a table entry and matches by name exactly as a live process does; the state column is the only field that separates them, and name-based tools do not print it.
CHECK Read the state rather than the count: `ps -o pid,stat,comm -p PID` — Z is dead, S, R or D are alive. Observed here: a forked child whose parent never waits shows STAT Z and `[python3] `, survives `kill -9` unchanged, is matched by `pgrep python3` and is not matched by `pgrep -f`, because /proc/PID/cmdline for a zombie is 0 bytes.
cost A shutdown loop waits forever on a process that has already exited, or a supervisor counting instances by name refuses to start the replacement it should have started.
generalises Every registry where deregistration is a third party's responsibility: service discovery entries, connection pool slots, lock rows, task records whose owner died.
source https://man7.org/linux/man-pages/man1/ps.1.html
provenance documented
NS-037 is-active reports active for a unit whose processes have all exited
reads as `systemctl is-active provisioning` prints active. Conclusion drawn: the provisioning service is running.
actually RemainAfterExit= 'specifies whether the service shall be considered active even when all its processes exited'. A Type=oneshot unit with RemainAfterExit=yes runs its command to completion and is then held active indefinitely with MainPID 0. Nothing is executing, and nothing will restart if the work it did is undone. Stock Ubuntu ships many such units — apparmor.service, cloud-config.service — all reading as 'active (exited)'.
blind because is-active collapses ActiveState to a single word. That word covers both a running process and a finished one-shot, and the field that distinguishes them, SubState, is not part of the answer.
CHECK `systemctl show -p SubState -p MainPID --value NAME` — `running` with a non-zero PID, or `exited` with MainPID 0. Observed here: a unit created with `systemd-run --user --property=Type=oneshot --property=RemainAfterExit=yes /bin/true` reports is-active `active`, SubState `exited`, MainPID `0` once /bin/true has returned.
cost A health check built on is-active passes for a service that finished minutes ago, and would keep passing if its binary were deleted afterwards. Restart= never fires either, because the unit is not running to fail.
generalises Every status vocabulary where one token spans 'in progress' and 'finished': job schedulers, CI stages, container states, queue workers reported as healthy.
source https://man7.org/linux/man-pages/man5/systemd.service.5.html
provenance documented
NS-038 A running process keeps executing a binary that has been replaced on disk
reads as `ps -o args= -p PID` shows the daemon running from /usr/local/bin/app, and /usr/local/bin/app contains the new build. Conclusion drawn: the new build is running.
actually Replacing an executable unlinks the old inode; a process that already mapped it holds a reference and goes on executing the previous image until it is restarted. ps prints the path recorded at exec time, which now names a different file. /proc/PID/exe still resolves, with the string '(deleted)' appended to the original pathname, and the mapped pages come from an inode with no directory entry left.
blind because ps prints a string, not an identity. The argv and the executable path are untouched by the upgrade, so the row is byte-identical before and after it.
CHECK Compare inodes rather than paths: `stat -Lc %i /proc/PID/exe` against `stat -c %i /path/to/binary`. Observed here: identical (1908455) before the upgrade; after `rm` and a fresh copy the file on disk was inode 1908456 while the process still resolved to 1908455, `readlink /proc/PID/exe` ended in '(deleted)' and /proc/PID/maps held five deleted entries. The ps output was the same in both cases.
cost A package upgrade or a deploy is verified against the file that was written while the old code keeps serving, until an unrelated restart changes the behaviour with no corresponding change to anything.
generalises Any consumer that resolves a resource once and holds it: loaded shared libraries, opened config files, cached DNS answers, imported modules.
mitigation `lsof +L1`, or `ls -l /proc/*/exe 2>/dev/null | grep deleted`, names every process still on an old image; needrestart does this after apt on Debian and Ubuntu.
source https://man7.org/linux/man-pages/man5/proc_pid_exe.5.html
provenance documented
NS-039 The process bearing the service's name is a wrapper, and the worker is its child
reads as ps shows one process named run-analytics; it is killed and it leaves the list. Conclusion drawn: the service is stopped.
actually A shell wrapper ending in a plain command rather than `exec` forks a child and waits on it, so there are two processes: the wrapper carrying the recognisable name, and the worker carrying the interpreter's or binary's own name. Killing the wrapper leaves the worker running, reparented to PID 1, still holding its port, its lock and its open files.
blind because The process list is flat and reports names. Nothing in it indicates that the row matching the search is the parent of the row doing the work, and after the kill the name is genuinely gone.
CHECK Verify the resource rather than the name: `ss -ltnp 'sport = :PORT'` after the stop — empty means stopped, a listener means the worker outlived the name. Observed here: killing `/bin/bash ./run-analytics` left its `sleep 120` child alive with PPID reassigned to 1; `pgrep -P PID` lists such children before the kill rather than after.
cost A restart yields two live workers competing for the same resource, or a stop that reports success while the old worker keeps writing. Neither produces an error.
generalises Every process tree observed as a flat list: container entrypoints, npm and make targets, virtualenv shims, ssh command wrappers.
mitigation `exec` as the last line of a wrapper replaces the shell instead of forking, so the name and the worker are one process; systemd's default control-group-based KillMode stops the whole tree rather than the named process.
source https://www.gnu.org/software/bash/manual/bash.html#Bourne-Shell-Builtins
provenance documented
NS-040 pgrep matches a fifteen-character truncation of the process name
reads as `pgrep analytics-ingest-worker` prints nothing and exits 1. Conclusion drawn: the process is not running, so it should be started.
actually The kernel stores a process name of at most sixteen bytes including the terminator, so ps and pgrep see 'analytics-inges'. The manual states it plainly: 'the process name used for matching is limited to the 15 characters present in the output of /proc/pid/stat'. Any pattern longer than that matches nothing, whatever is running.
blind because The tool returns a true statement about a name it truncated, and an exit status identical to the one for 'no such process'. Newer procps prints a warning, but on stderr — the stream discarded by exactly the scripts that make this mistake.
CHECK Match the command line instead: `pgrep -af analytics-ingest-worker`. Observed here with `./analytics-ingest-worker 60` running: `pgrep analytics-ingest-worker` printed nothing and exited 1, /proc/PID/comm contained 'analytics-inges', and both `pgrep -f analytics-ingest-worker` and `pgrep analytics-inges` found the process.
cost A guard that starts the service when the check fails starts a second copy of something already running: two writers on one database, two schedulers on one queue, each verified as absent immediately beforehand.
generalises Every lookup key silently normalised before comparison: truncated identifiers, case-folded names, unicode-normalised paths, hostnames cut at a label boundary.
source https://man7.org/linux/man-pages/man1/pgrep.1.html
provenance documented
NS-041 A process list taken in one namespace describes a different machine
reads as `ps aux` inside the container lists the worker as PID 1 and little else. Conclusion drawn: this is what is running on the machine, and PID 1 is the thing to signal.
actually A PID namespace isolates a set of process IDs: a process has a different PID in each namespace it belongs to, and processes outside the namespace are invisible from within it. The listing enumerates one view. The same worker holds another PID on the host, every host process competing for the same CPU and memory is absent, and a PID copied from one view and used in the other addresses an unrelated process, if it addresses anything.
blind because PID numbers are namespace-relative and are printed as bare integers with nothing to record which namespace produced them. Two listings of 'the processes' are simply two different sets, each internally consistent.
CHECK Compare the observer's namespace with that of the process being acted on before trusting the number: `readlink /proc/self/ns/pid` against `readlink /proc/PID/ns/pid`. Identical inode strings mean the PIDs are comparable; different ones mean they are not. Observed on this host both returned `pid:[4026531836]` and `systemd-detect-virt --container` returned `none` — the case the check exists to establish rather than assume.
cost Resource accounting done inside a container attributes to itself memory the host is losing elsewhere, and a kill or restart aimed at a PID from the other view lands on whatever holds that number there.
generalises Every identifier unique only within a scope the output does not name: container PIDs, per-tenant row ids, per-session handles, relative paths.
source https://man7.org/linux/man-pages/man7/pid_namespaces.7.html
provenance documented
NS-063 free and ps report the host's memory, not the limit the process runs under
reads as `free -h` inside the workload reports 7.8 GiB total and 4.1 GiB available. Conclusion drawn: memory is plentiful, so a slowdown or a death has some other cause.
actually The limit is a cgroup property and /proc/meminfo is not scoped to it. memory.max is 'the main mechanism to limit memory usage of a cgroup. If a cgroup's memory usage reaches this limit and can't be reduced, the OOM killer is invoked in the cgroup.' The process may be reclaiming continuously against a ceiling two orders of magnitude below the figure free prints.
blind because free, top and ps read /proc, which describes the machine. The constraint lives in /sys/fs/cgroup, which they do not consult.
CHECK Read the limit and the pressure counters for the process's own cgroup: `CG=$(awk -F: '{print $3}' /proc/self/cgroup); cat /sys/fs/cgroup$CG/memory.max /sys/fs/cgroup$CG/memory.events`. Observed on this host inside `systemd-run --user --scope -p MemoryMax=200M`: free -h still reported 7.8Gi total and 4.1Gi available, memory.max read 209715200, and a 400 MB allocation reported success while memory.events moved from `max 0` to `max 772`, recording 772 occasions on which the limit was hit and reclaim forced.
cost Tuning is done against the wrong ceiling. A process being throttled or killed by a limit is diagnosed as slow code, and the counter that would have said so was never read.
generalises Every constraint enforced at a layer the inspection tool does not model: cgroup CPU quota against nproc, container disk quotas against df, API rate limits against local concurrency settings.
source https://docs.kernel.org/admin-guide/cgroup-v2.html
provenance documented
NS-064 Load average counts processes blocked on disk as though they were running
reads as Load average is 8.0 on a four-core machine. Conclusion drawn: the CPU is saturated, so the answer is fewer workers or more cores.
actually proc_loadavg(5): the first three fields give 'the number of jobs in the run queue (state R) or waiting for disk I/O (state D)'. Uninterruptible sleep is counted the same as running. A load of 8 alongside an idle CPU means eight processes are blocked on storage or a stalled network filesystem, and adding cores changes nothing.
blind because The load average is a single number covering two unlike conditions. Nothing in it separates work being done from work waiting to begin.
CHECK Count the states behind the number: `ps -eo state= | sort | uniq -c`. Observed on this host at load 2.14: 101 processes in S, 74 in I, 2 in R and none in D, so the load is runnable work rather than blocked I/O. Under an I/O stall the same command shows the D column carrying the figure.
cost Capacity is added to a machine that is not short of capacity, the stall persists, and the real cause, a slow volume or a wedged mount, is never examined.
generalises Every aggregate that sums dissimilar states: queue depth mixing retries with new work, request counts mixing served with rejected, connection counts including half-open ones.
source https://man7.org/linux/man-pages/man5/proc_loadavg.5.html
provenance documented
NS-065 The %CPU column is a lifetime average, not a current rate
reads as `ps aux` shows one process at 85.1% CPU. Conclusion drawn: this is the process consuming the machine now.
actually ps(1) states it plainly: 'CPU usage is currently expressed as the percentage of time spent running during the entire lifetime of a process.' A process that pinned a core for six seconds and has been idle since still reports a high figure, decaying only as its lifetime grows. The converse matters more: a process running for a day that began spinning a minute ago reports a small number.
blind because The column has the units of a rate and is read as one. Nothing in the output records the averaging window, which is each process's own age and therefore different in every row.
CHECK Measure the delta over a known interval: read fields 14 and 15 of /proc/PID/stat twice and divide the difference by CLK_TCK times the elapsed seconds. Observed on this host with a process that spun for six seconds and then slept: ps reported 85.1% immediately afterwards and 22.0% twenty seconds later, while the tick delta over the following three seconds was 0 out of 300 possible, that is 0.0% actual.
cost The wrong process is restarted, throttled or blamed, and the one that has quietly started to spin is ranked below it because its long life dilutes its average.
generalises Every statistic whose window is implicit: lifetime averages, cumulative counters presented as gauges, uptime-normalised error rates.
mitigation top's %CPU is an interval rate rather than a lifetime average, and pidstat reports per-interval figures directly.
source https://man7.org/linux/man-pages/man1/ps.1.html
provenance documented
NS-086 kill reports success when the signal was delivered and disregarded
reads as `kill $PID` exits zero and the deploy script moves on. Conclusion drawn: the old worker has stopped.
actually kill(2): on success, at least one signal was sent, zero is returned. Success means the signal was queued to a process the caller had permission to signal. A process that installed an ignore disposition for SIGTERM, whether through `trap '' TERM` or a runtime that swallows it while a shutdown hook stalls, receives the signal and carries on. signal(7) notes the only exceptions: SIGKILL and SIGSTOP cannot be caught, blocked, or ignored.
blind because The status describes the sender's half of the transaction. Whether the recipient acted is a fact about the recipient, observable only afterwards and only by looking again.
CHECK Read the target's signal dispositions, or simply look again after a pause: `grep -E '^Sig(Ign|Blk|Cgt)' /proc/$PID/status`. Observed on Linux 6.8 with a script carrying `trap '' TERM` and `trap '' HUP`: two successive `kill` invocations both exited 0 and the process was still listed by `ps` after each, reporting `SigIgn: 0000000000004005`, the bits for signals 1, 3 and 15. `kill -9` ended it. `os.kill` against an unreaped zombie likewise raised nothing and returned normally.
cost The deploy continues believing the port is free. The replacement either fails to bind, or binds elsewhere and serves alongside the process that was supposed to be gone, producing a fleet where half the requests run old code.
generalises Every asynchronous request whose acknowledgement is acceptance of the message rather than performance of the work: signals, queue publishes, webhook deliveries, cache invalidations.
mitigation Treat termination as a condition to be waited on rather than an instruction to be issued: poll for the process to disappear, with a bounded escalation to SIGKILL.
source https://man7.org/linux/man-pages/man2/kill.2.html
provenance documented
NS-087 A unit reported inactive can still have every worker it started running
reads as `systemctl stop app` exits zero and `systemctl is-active app` prints inactive. Conclusion drawn: the service and everything it spawned are stopped.
actually With KillMode=process, systemd.kill(5) states that only the main process itself is killed (not recommended!), and warns that this allows processes to escape the service manager's lifecycle and resource management, and to remain running even while their service is considered stopped and is assumed to not consume any resources. The workers keep their sockets, locks and memory. The default, control-group, kills the whole cgroup and does not have this behaviour.
blind because is-active reports the unit's state, and the unit's state is decided by its main process. Once the survivors have outlived the unit they are no longer accounted to it, so the supervisor's view is accurate and incomplete at once.
CHECK Ask the kernel who is alive rather than asking systemd whether the unit is: `ps -eo pid,ppid,args | grep '[w]orker'`, or `ss -ltnp` for the port the service held. Observed on systemd 255 with a user unit `Type=simple` and `KillMode=process` whose ExecStart backgrounded a child: `systemctl --user stop` exited 0, `is-active` printed inactive, and `ps` still listed the child at pid 3917771. The identical unit at the default KillMode=control-group left nothing behind.
cost A restart appears to work while the previous generation continues serving; two versions run concurrently and diverge, and the resources the orphans hold are invisible to anything that accounts by unit.
generalises Any supervisor whose notion of the service is narrower than the set of processes the service created: init systems, container runtimes, CI job runners, test harnesses spawning fixtures.
mitigation Leave KillMode at control-group unless there is a specific reason not to, and verify a stop by the absence of processes and listeners rather than by the unit's state.
source https://www.freedesktop.org/software/systemd/man/latest/systemd.kill.html#KillMode=
provenance documented
======================================================================
INSTRUMENT: log-output — Logs and stdout
======================================================================
Used when: You read the output and it looked normal.
Captures: What the program chose to say about itself, in the branches that say anything.
Cannot see:
- Silent branches, which by construction produce no line to read.
- Failures attributed to the wrong actor, where the reported victim is not the cause.
- Anything killed before it could flush a buffer.
NS-011 The OOM killer names the fattest process, not the one that leaked
reads as A long-running session dies mid-task. The log records that session being killed. Conclusion drawn: that session was the problem.
actually A different process had leaked for hours — 193 browser instances spawned by automation and never closed, 27 still resident. The OOM killer selects by current footprint, so it shot the largest process, which was an unrelated session whose context had simply grown. The leak and the casualty were different processes.
blind because The log faithfully records the victim. It has no field for the cause, and nothing in the kill message distinguishes 'grew large' from 'made the machine run out'.
CHECK Rank every process by RSS at the time of death, not just the one named: ps -eo rss,comm --sort=-rss | head -20, and count instances of anything spawned in a loop. A single fat process is a victim; a hundred medium ones are the cause.
cost The innocent session is blamed and 'fixed'. The leak keeps running and takes another process later.
generalises Every resource-exhaustion system that reports which tenant it evicted rather than which one filled the resource.
mitigation Cap or close anything spawned per-iteration, and check free memory before adding load rather than after losing work.
provenance observed
NS-028 A rotated log leaves the daemon writing to a file that no longer has a name
reads as app.log exists, is zero bytes, and gains no lines. Conclusion drawn: the service is idle, or has stopped working.
actually logrotate renamed or removed the file the process had open. The process still holds the old inode and keeps appending to it. The logrotate man page names the case in its description of copytruncate: it exists for programs that cannot be told to close their logfile and thus might continue writing to the previous log file forever.
blind because Reading a log means resolving a path. After rotation the path and the process's open descriptor refer to different objects, and the reader follows the path while the writer holds the descriptor.
CHECK Ask the process which file it is writing to: `ls -l /proc/$(pidof app)/fd | grep -i log`. A healthy process points at the live path; a stranded one points at a path marked `(deleted)`.
cost Log-based monitoring goes quiet and the quiet is read as calm. Disk fills with a file no directory listing can show, and it is only reclaimed when the process is restarted.
generalises Any handle held across a rename or delete: log files, config files watched by path, unlinked sockets and temp files.
mitigation copytruncate, or a postrotate hook that signals the daemon to reopen its log.
source https://man7.org/linux/man-pages/man8/logrotate.8.html
provenance documented
NS-029 Python discards records below WARNING when no logging is configured
reads as A script instrumented with logger.info() at every step produces no output at all. Conclusion drawn: the code path never ran.
actually With no configuration, the root logger has no handlers and the internal last-resort handler is set at WARNING. INFO and DEBUG records are created and then dropped; WARNING and above go to stderr. The code ran, and said so, into nothing.
blind because A discarded record and a record that was never emitted produce the same empty output. A log cannot report what it filtered out, because the filtering happens before anything is written.
CHECK `logging.getLogger(__name__).isEnabledFor(logging.INFO)` — False while records are being dropped, True once a handler and level are configured. Observed on 3.12: root handlers `[]`, lastResort `<_StderrHandler (WARNING)>`, isEnabledFor(INFO) False.
cost Debugging proceeds from the false premise that the instrumented branch was not reached, and the real fault is hunted upstream of where it lives.
generalises Every level-filtered or sampled telemetry channel, where the absence of a line is read as the absence of an event.
source https://docs.python.org/3/howto/logging.html#what-happens-if-no-configuration-is-provided
provenance documented
NS-031 A killed process loses the output it produced but never flushed
reads as A worker is killed and its log ends several steps before the operation under investigation. Conclusion drawn: execution never reached that step.
actually Standard output is block-buffered whenever it does not refer to a terminal, so lines accumulate in a user-space buffer of a few kilobytes until it fills or the process exits cleanly. SIGKILL cannot be caught, blocked or handled, so no flush happens. The steps ran, announced themselves, and the announcements died in the buffer.
blind because A log records what was flushed, not what was written. A line that was never produced and a line that was produced into a buffer and then discarded are the same absence.
CHECK Re-run with buffering removed and kill it the same way: `PYTHONUNBUFFERED=1 prog > out.log` (or `stdbuf -oL` for a C program). Observed on 3.12: a script printing two lines and then sleeping, SIGKILLed two seconds in, left out.log at 0 bytes; the identical run under PYTHONUNBUFFERED=1 left both lines. Output that appears only when unbuffered was being produced all along.
cost The investigation moves upstream of the last logged line, which is not where the process was. The fault lives inside the region the log appears to prove was never entered.
generalises Every buffered channel inspected after an abrupt stop: stdio, log shippers with in-memory queues, metrics flushed on a timer, traces batched before export.
mitigation Anything whose log will be read after an abnormal death should be unbuffered at the point of writing; adding it at the point of reading is too late.
source https://docs.python.org/3/library/sys.html#sys.stdout
provenance documented
NS-032 journald discards every message past the burst and files the notice elsewhere
reads as `journalctl -u worker` covers the whole run and contains no error and no completion line. Conclusion drawn: the worker raised no error, and the absent completion line is the anomaly worth chasing.
actually If more messages than RateLimitBurst are logged by a service inside RateLimitIntervalSec, all further messages within the interval are dropped until the interval is over. The default is 10000 messages in 30s, multiplied by a factor derived from the free disk space available to the journal. Everything the service says after the burst is exhausted is discarded, including the line that mattered.
blind because The stored records are contiguous and well-formed; the log simply stops and later resumes. A message about the number of dropped messages is generated, but by journald under its own identity, so a query filtered to the unit does not show it — and a reader outside the systemd-journal and adm groups cannot see it at all.
CHECK Count what the producer emitted against what the journal stored. Observed on this host: a transient unit emitting 120,001 numbered lines in 2.1s stored 37,499 of them — line 1 through line 37,499 and then nothing at all, with the final line absent and no suppression notice visible under `journalctl --user -u NAME`.
cost A verbose service is treated as a well-instrumented one, and its silence during the interesting minute is read as calm rather than as the direct consequence of its own verbosity.
generalises Every sampled or throttled telemetry path: metrics agents, trace sampling, syslog rate limits, ingestion quotas in hosted log services.
mitigation LogRateLimitIntervalSec= and LogRateLimitBurst= can be raised per unit, but a service logging at that rate needs to log less rather than louder.
source https://man7.org/linux/man-pages/man5/journald.conf.5.html
provenance documented
NS-033 basicConfig does nothing once anything has already touched the root logger
reads as `logging.basicConfig(level=logging.DEBUG)` runs at the top of main() and the program still emits only warnings. Conclusion drawn: the instrumented branches are not being reached, or the level argument is wrong.
actually The function does nothing if the root logger already has handlers configured, unless force is set to True. One earlier call to logging.warning(), one imported library that logs during import, or one framework that configures logging first, installs a handler; every later basicConfig call is then a no-op and the root level stays at WARNING.
blind because The call raises nothing and returns nothing to inspect, and the handler that is present is a working handler emitting real lines. The output is a correct log at the wrong level, which is far more convincing than no log at all.
CHECK Interrogate the configuration rather than the output: `python3 -c 'import logging; logging.warning("x"); logging.basicConfig(level=logging.DEBUG); print(logging.root.handlers, logging.root.level, logging.getLogger().isEnabledFor(logging.INFO))'`. Observed on 3.12: handlers `[ (NOTSET)>]`, level 30, isEnabledFor(INFO) False; with `force=True` the INFO record appears and isEnabledFor(INFO) is True.
cost Missing lines are attributed to unreached code, and the debugging effort goes into the application instead of into the two lines of logging setup that declined to apply.
generalises Every initialiser that is idempotent by doing nothing: first-wins registries, singleton bootstrappers, setup functions that check for prior state and return quietly.
source https://docs.python.org/3/library/logging.html#logging.basicConfig
provenance documented
NS-034 A malformed log call discards its own record and returns normally
reads as app.log holds the lines either side of a payment and no line for the payment itself. Conclusion drawn: that branch did not execute.
actually The argument count did not match the format string. Interpolation happens inside the handler rather than at the call site, so the exception is raised during emit() and routed to Handler.handleError. The call returns normally, the program continues, the exit status is 0, and the record is gone. Where logging.raiseExceptions has been set to False — 'this is what is mostly wanted for a logging system' — nothing is written anywhere.
blind because The line is missing for a reason internal to the logging subsystem, and the logging subsystem is the instrument being read. Its own failures are the one class of event it is built not to report through itself.
CHECK Look on the other stream, which is where the default handler puts its own failures: `python3 app.py 2>&1 >/dev/null | grep -c '^--- Logging error ---'` — non-zero when records were formatted and thrown away, zero when the branch genuinely did not run. Observed on 3.12: `log.info('charged %s for %s', 'user-1')` produced a TypeError traceback on stderr, exit status 0, and an app.log containing only the following line; with `logging.raiseExceptions = False` stderr was 0 bytes and the log was identical.
cost The audit trail has a hole exactly where an operation is hardest to reconstruct, and the hole is read as evidence the operation did not happen.
generalises Every subsystem asked to report on itself: monitoring agents that cannot alert on their own death, error trackers that drop malformed events, audit logs that fail open.
mitigation Route stderr to the same destination as the log, so the logging subsystem's own failures land beside the records they replaced.
source https://docs.python.org/3/library/logging.html#logging.Handler.handleError
provenance documented
NS-035 Combined stdout and stderr arrive in an order that never happened
reads as `prog > run.log 2>&1` yields three failure lines followed by three start lines. Conclusion drawn: the failures preceded the work, so something failed before the steps began.
actually Standard output is fully buffered when it does not refer to an interactive device; standard error is not. Redirected to a file, stdout accumulates and is written in one block at exit while every stderr line goes straight through. The order in the merged file is an artefact of buffering policy, not a record of time.
blind because A log is read as a sequence and a sequence is read as causality. The merged file records neither the originating stream nor the moment of emission, so the interleaving is the only ordering evidence available and it is precisely the part that was destroyed.
CHECK Re-run with stdout unbuffered and compare the two files. Observed on 3.12: a program alternating a stdout line and a stderr line three times produced all three stderr lines before all three stdout lines under `> combined.log 2>&1`, and strictly alternating lines under `PYTHONUNBUFFERED=1`. `stdbuf -oL` did not change it, because the interpreter manages its own buffers rather than libc's.
cost A cause is assigned to the wrong step, and the fix is applied to whatever the reordering happened to place first.
generalises Any merge of independently buffered sources into one ordered view: multi-process logs, distributed traces without synchronised clocks, tail -f across several files.
mitigation Timestamp at the point of emission, so ordering does not depend on arrival, and keep the two streams separate when their relative order carries meaning.
source https://man7.org/linux/man-pages/man3/stdio.3.html
provenance documented
NS-066 Redirection order decides whether the log can contain errors at all
reads as The service runs as `app 2>&1 > app.log` and app.log holds a clean sequence of startup lines with no errors. Conclusion drawn: the run was clean.
actually Redirections are processed left to right. The bash manual gives this exact pair: `ls > dirlist 2>&1` sends both streams to the file, while `ls 2>&1 > dirlist` 'directs only the standard output to file dirlist, because the standard error was duplicated from the standard output before the standard output was redirected to dirlist'. Standard error went wherever standard output pointed beforehand, usually a terminal that no longer exists or a parent's discarded output.
blind because The log is genuine, complete and correctly ordered for the stream it captured. Nothing in it can indicate that a second stream existed and went elsewhere, so a log with no errors and a log that cannot contain errors are the same file.
CHECK Ask the running process where its descriptors point: `readlink /proc/$$/fd/1 /proc/$$/fd/2` from inside the redirected command. Observed on bash 5.2.21: under `./probe.sh 2>&1 > out1.log`, fd 1 pointed at out1.log while fd 2 pointed at the parent's output; under `./probe.sh > out2.log 2>&1` both pointed at out2.log. A script emitting one error line produced `grep -c ERROR` of 0 in the first case and 1 in the second.
cost Every diagnostic the program emits is discarded by the same arrangement that produces the record used to declare it healthy, and alerting built on that log's contents can never fire.
generalises Any capture configured to watch one channel while the interesting events use another: stderr against stdout, structured logs against panics, application logs against the supervisor's.
source https://www.gnu.org/software/bash/manual/bash.html#Redirections
provenance documented
NS-067 A service's stdout is recorded at info priority whatever the line says
reads as `journalctl -u app -p err` prints '-- No entries --'. Conclusion drawn: the service has logged no errors.
actually systemd assigns the priority, not the text. SyslogLevel= is 'the default syslog log level to use when logging to the logging system or the kernel log buffer', it 'only applies to log messages written to stdout or stderr', and it 'Defaults to info'. Unless a line carries an explicit angle-bracket level prefix, every line the process prints is stored at priority 6, including the ones whose text reads ERROR.
blind because The filter and the store agree. Priority is metadata attached at ingestion, so a severity filter reports the transport's opinion rather than the application's.
CHECK Look at the priority distribution instead of the filtered view: `journalctl -u UNIT -o json | jq -r .PRIORITY | sort | uniq -c`. Observed on a unit logging 43,061 records over three weeks: every one at PRIORITY 6 (informational), while a plain-text search of the same range found 14 lines containing 'error'. `journalctl -u UNIT -p err` reported '-- No entries --' throughout.
cost Severity-based alerting and triage are silently disabled for every service that logs to stdout without prefixes, which is most of them, and the absence of high-priority records is read as the absence of high-priority events.
generalises Every field assigned by a collector rather than by the source: levels inferred by a log shipper, statuses rewritten by a proxy, timestamps stamped at ingestion.
mitigation Emit the angle-bracket level prefix from the application, or set SyslogLevel= on the unit; SyslogLevelPrefix= controls whether such prefixes are honoured.
source https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#SyslogLevel=
provenance documented
NS-068 A journal with no persistent directory discards its evidence at reboot
reads as After a crash and a restart, `journalctl -u app --since '2 days ago'` returns nothing. Conclusion drawn: the service logged nothing before it died, so the failure was abrupt.
actually journald's Storage= defaults to auto, and 'auto behaves like persistent if the /var/log/journal directory exists, and volatile otherwise (the existence of the directory controls the storage mode)'. Under volatile storage the journal lives below /run and does not survive a reboot. The pre-crash records existed, and the restart performed to recover deleted them.
blind because An empty query result has one shape. 'Nothing was logged', 'nothing matched the filter' and 'the storage that held it no longer exists' all render as no output and a zero exit.
CHECK Establish whether history survives before drawing conclusions from its absence: `ls -d /var/log/journal 2>/dev/null; journalctl --list-boots`. Observed on this host: /var/log/journal exists and holds 566 MB, and --list-boots lists two boots reaching back five weeks, so an empty result here is a fact about the service. On a host without that directory the same commands print nothing and a single boot, and no empty result carries information.
cost The post-mortem proceeds from the premise that the process died silently, and the reboot performed to restore service is the act that destroyed the evidence for any other explanation.
generalises Every store whose retention is shorter than the investigation: ring buffers, in-memory metrics, container logs removed with the container, tmpfs working directories.
mitigation Create /var/log/journal and restart systemd-journald, or set Storage=persistent explicitly, before the next incident rather than after it.
source https://www.freedesktop.org/software/systemd/man/latest/journald.conf.html#Storage=
provenance documented
NS-069 Traffic classified by user-agent counts what clients claim to be
reads as An access log shows 57% of requests from browsers. Conclusion drawn: most visitors are people, and the site is reaching a human audience.
actually The User-Agent header is set by the client and asserted, never verified. A single scanner sending a stock Windows Chrome string produced a large share of that bucket; its requests were for /contact, /about-us, /pricing, /team and /support — pages this site has never had. Real browsers and anything imitating one are indistinguishable by header alone.
blind because The field being counted is supplied by the party being measured. Every row is internally consistent and none of them is evidence.
CHECK Group requests by client address and compare what each one asked for against what exists. A client whose requests are mostly 404s for pages the site has never published is enumerating, whatever it calls itself. Corroborate with an independent signal the client does not control, such as whether it also fetched the page's own subresources.
cost Audience is misread in the direction that flatters. Content decisions get made for readers who were never there, and genuine machine traffic is filed as human.
generalises Any metric derived from a field the measured party supplies: referrers, self-reported versions, declared content types, client-side analytics events.
mitigation Treat the user-agent as one weak signal among several. Behaviour — which paths, in what order, with which subresources — is set by the client too, but it is far more expensive to fake convincingly.
provenance observed
NS-088 Records longer than a pipe's atomic limit are spliced into one another
reads as `grep 'request_id=abc123' app.log` returns nothing, and the file around that period is full of well-formed lines. Conclusion drawn: that request never reached this service.
actually pipe(7): POSIX.1 says that writes of less than PIPE_BUF bytes must be atomic, the output data being written to the pipe as a contiguous sequence, while writes of more than PIPE_BUF bytes may be nonatomic, and the kernel may interleave the data with data written by other processes. On Linux PIPE_BUF is 4096 bytes. Several workers writing to one pipe, which is what a container's stdout, a `tee` and most log shippers are, produce records cut open with another worker's record inserted into the gap. The result still ends in a newline, so it is still a line.
blind because A log reader sees lines, and a spliced line is a line: it has a beginning, an end and plausible contents. The pattern that would have matched now straddles a boundary that did not exist when the record was written, and the line count is unchanged.
CHECK Validate each line against the format the writer emits and count the failures, rather than counting lines. Observed on Linux 6.8 with four writers into one pipe behind a deliberately slow reader: at 4090-byte records, 800 lines and 0 malformed; at 5000-byte records, 800 lines and 21 malformed; at 20000-byte records, 800 lines and 195 malformed, one of which opened with `BEGIN-B-0010` and contained an entire `BEGIN-A-0000 ... END-A-0000` record inside it. The line count was 800 in every run.
cost Requests appear never to have happened, error rates read low, and the records that would contradict both are present in the file in a form no query will match.
generalises Any shared append-only channel with an atomicity limit: pipes, datagram sockets, unlocked file writes, records assembled from several write calls.
mitigation Keep each record under PIPE_BUF, or give each writer its own descriptor opened O_APPEND onto a regular file, where appends do not interleave regardless of size.
source https://man7.org/linux/man-pages/man7/pipe.7.html
provenance documented
NS-089 A log search returns nothing because the window and the timestamps are in different zones
reads as `journalctl -u app --since '2026-08-24 00:20:00'` prints `-- No entries --` for a window that covers the incident. Conclusion drawn: the service logged nothing then, so it was not running or was never reached.
actually journalctl interprets --since and --until in the local time zone, and systemd.time(7) states that on display systemd will format timestamps in the local timezone. When the window is copied from a source in another zone, a UTC dashboard, a cloud console, an API response or a colleague on another continent, the query addresses a moment hours away from the one intended. The entries exist and sit outside the range.
blind because An empty result set has one shape. Nothing separates 'no entries in this window' from 'the window was somewhere else', and the timestamps that would reveal the offset are precisely the ones the filter excluded.
CHECK Ask for the entries in an unambiguous frame and see whether they exist at all before filtering: `journalctl -u app -n 5 --utc -o short-iso`. Observed on this box (Etc/UTC) against a single `logger -t vftz` entry: plain `journalctl -t vftz` displayed it as `Aug 24 00:24:27`, while `TZ=America/New_York journalctl -t vftz` displayed the same entry as `Aug 23 20:24:27`, a different calendar day. Passing a window taken from the UTC clock while TZ was America/New_York returned `-- No entries --` for a record written seconds earlier.
cost The investigation concludes the service was silent during the incident and moves upstream, while the evidence sits in the same file a few hours away.
generalises Every filter expressed in units the store does not share: time zones, seconds against milliseconds, inclusive against exclusive bounds, severities named differently by the writer and the query.
mitigation Pin both ends of every correlation to one frame: query with --utc and read with -o short-iso, or attach an explicit offset to every timestamp that crosses a system boundary.
source https://www.freedesktop.org/software/systemd/man/latest/journalctl.html
provenance documented
NS-090 Crawler hits in an access log are not evidence that a page is indexed
reads as The access log shows repeated fetches from Googlebot, Bingbot and other declared crawlers, and the sitemap was accepted. Conclusion drawn: the pages are in the index and the site is discoverable.
actually Crawling, indexing and ranking are three separate stages. A crawler fetching a URL records only that it was retrieved; the page may then be excluded, deduplicated against similar content, or held in a queue for days. A site can be fetched hundreds of times and return no results for a search of its own exact title.
blind because The access log is written by the origin and can only record requests that reached it. It has no field for what the requester did afterwards, and no stage of indexing produces a request back to the server.
CHECK Query the index itself rather than reading the log: search for an exact phrase unique to the page, in quotes, and separately run a `site:` query for the domain. Both return nothing while the page is merely crawled. For a property you control, the index-coverage report in Google Search Console or Bing Webmaster Tools states the stage per URL.
cost Distribution work is reported as finished on the strength of crawler traffic, and the weeks in which the pages are fetched but unfindable pass unnoticed. Effort moves on to new content while nothing published so far can be reached by search.
generalises Any pipeline whose early stages report back to you and whose later stages do not: submitted-versus-accepted, queued-versus-delivered, uploaded-versus-published, deployed-versus-serving.
mitigation Treat submission and crawling as inputs, not outcomes. The observable outcome is a result page containing your URL.
provenance observed
======================================================================
WHAT THIS CANNOT SEE
======================================================================
This is a reference about instruments that cannot see their own blind spots. It would be a poor one if it did not state its own.
Out of scope:
- Failures that announce themselves: A crash, an exception, a 500, a red test — these are cheap. They tell you where to look. Everything here is chosen because it does not.
- Security vulnerabilities: A different discipline with its own literature and threat models. Some entries touch it — bidirectional control characters, credentials dropped across a redirect — but only where the failure is silent, never as security coverage.
- Whether the requirement was right: Every check here answers 'did this do what you asked'. None answer 'was asking for it correct'. No instrument sees a wrong specification perfectly implemented.
- Performance work: Except where slowness is misattributed — an asset queued rather than slow, a lifetime CPU average read as current load. Tuning is not the subject.
- Flaky failures: Something that fails one run in fifty is visible, just rarely. The entries here fail every time and look correct every time.
Known bias:
- Unix, the web, and interpreted languages: Of 74 cited sources, 19 are Linux man pages and 11 are RFCs. There is nothing here from Windows, macOS, mobile, embedded, or the JVM and .NET ecosystems. Databases appear once. Machine-learning pipelines, message queues and distributed consensus do not appear at all. Those platforms have silent failures; this registry has not looked at them.
- Verified on one machine: Checks were reproduced on Linux with Python 3.12, bash 5.2, systemd 255, curl 8.5 and Chrome 151. Behaviour differs across versions, and some entries record a behaviour that a later release may fix. Where a version mattered it is named in the entry.
- Written quickly, by few authors: Most entries were authored in a single day. A reference like this should accrete from things that actually went wrong, over months. Read it as a strong start rather than a settled body of knowledge.
- Only failures somebody noticed: This is the deepest limit and it cannot be fixed from inside. An entry exists because a person or an agent eventually caught the failure and worked out the mechanism. Failures that are silent AND have never been caught are, by construction, absent — and there is no way to estimate how many there are. The registry's own instrument is 'somebody noticed', and it is blind to exactly what it is about.
======================================================================
PRINCIPLES
======================================================================
P-1 Prefer the resolved value over the authored one.
Configuration files record intent. Computed styles, running processes and served bytes record outcome. When they disagree, only one of them is what users experience.
P-2 A check is only diagnostic if it can come out either way.
An observation that returns the same result under both hypotheses has confirmed nothing, however much work it took to produce.
P-3 Capture errors before adjusting values.
Inert and wrong look identical from the outside. One error capture distinguishes them; no amount of parameter tuning does.
P-4 Know your instrument's failure modes before trusting its readings.
A screenshot cannot see time. An exit code cannot see semantics. A cache cannot see freshness. Each is silent about exactly the thing it cannot represent.
P-5 Distrust the fallback that is good enough.
Degradation designed to be invisible to users is equally invisible to the agent verifying the work.
P-6 Report the observation, not the inference.
'The service is active and returned OK' can be verified by a reader. 'It works' cannot.
P-7 Silence is not evidence of success.
A process that hangs, a branch that logs nothing, and a command that was never reached all produce the same empty output as a clean run.
P-8 Check whether you are inside what you are measuring.
Searches that match themselves and teardowns that destroy their own host are the same error: the observer was part of the sample.
Machine-readable: https://verifyfirst.dev/registry.json