Evidence & Inquiry

The defect where everything passes

A defect family that produces well-formed output from wrong or absent input and raises nothing at all. Four places it hides, one rule that finds it, and why fluent output makes the defect easy to miss.

What must you check when nothing is failing?

There is a kind of defect that no ordinary check will find, because every ordinary check passes. It has one signature, and the signature is short:

wrong input, well-formed output, no error.

Nothing throws. Nothing logs. The result arrives on time, in the right shape, with the right fields populated and the right types, and it is wrong. The fluency is part of the failure rather than a detail beside it: a result that arrives clean is a result nobody has a reason to open.

This is a method, not a result. It came out of a long audit whose only real discipline was that it had to be able to refute its author’s own prior conclusions, and it survives leaving that audit because none of it depends on the system it was built against.

Why a passing check is weak evidence

A test asserts that something is true. It does not assert that the thing it tested is the thing you care about.

Consider the difference between these two statements:

  • No error was raised.
  • The right thing happened.

Most verification I have read treats the first as evidence for the second. It is not. An error is raised only where somebody, in advance, imagined this particular way of being wrong and wrote the raise. The whole value of a defect in this family is that nobody imagined it, which is precisely why there is no raise to catch.

So the question that finds these is not what is failing. It is: for every branch, default, guard and constructor on this path, what happens when its input is absent or unrecognised? If the honest answer is “something well-formed comes out”, you have a finding, whether or not a test currently fails.

Four places it hides

Each of these four was read back out of a confirmed instance: the hit came first, the general shape afterwards. That order is the whole guarantee. A probe invented in advance describes the author’s imagination; a probe derived from something that actually happened describes the world.

One. A branch selected by name rather than by what the artifact declares. A component chooses how to treat an input by matching a string - a filename, a family name, a label someone typed into a configuration. The input carries authoritative metadata that answers the same question exactly, and the code does not read it. When the string does not match, the input takes a branch meant for something else entirely. It is handled. It is handled wrongly, in silence, and the output is perfectly well-formed.

Two. A default that is the rare case. Every dispatch needs a fallback, and the fallback is usually written early, when only one case exists. Later the common case arrives and the fallback is never revisited. Now anything unrecognised silently takes the exotic path, which is the opposite of what a default is for. The test for this is a counting question, not a logic question: of the inputs this system actually receives, how many does the fallback serve correctly? If the answer is none, the default is inverted.

Three. A compiled path with no coverage, behind a guard that is correct. A specialised implementation is written, guarded properly, and never executes, because an earlier branch always wins on every machine anybody runs. The guard is not the defect. The guard is why nobody noticed the code is unexercised, and unexercised code is not correct code - it is code with no evidence either way. The question is not “is this guarded” but “on the hardware we actually own, has this line ever run?”

Four. A structure constructed empty and served as though it were seeded. This is the sharpest of the four, and the most embarrassing to find in your own work. A constructor builds the right shape with nothing in it. A consumer receives it, does exactly what it is supposed to do, and returns output of the right dimensions that ignores the input entirely. Every interface contract is honoured. The shape is the alibi.

Notice what the four have in common. In none of them is the code wrong in the way a reviewer looks for wrongness. In all four, something was shaped correctly and empty of the thing that gives it meaning.

The artifact is the evidence; the text about it is a claim

The second half of the method is duller and does more work.

A comment, a filename, a configuration key and a document all describe the artifact. None of them is the artifact. Where the description and the artifact can disagree, the artifact is the evidence and the description is a claim about it.

In practice this turns into a table you can write for any system, with the convenient property that the right-hand column is always cheap:

What was claimed What was actually consulted
A configuration listing the available components The directory those components are read from
A document saying a format is supported The format identifier inside the file itself
A declared vocabulary size Whether the structure holding it is ever written to
“The model is served” Output from two deliberately different inputs
“This field is required” Whether the field exists in the artifact at all

Every row of that table is a check that costs one command. Every row of that table, in the audit it came from, changed a conclusion that was already written down.

The failure mode of skipping it is not an error. It is a document that reads as verified and is not - which is the same defect family, one level up. Prose is also a well-formed output.

The discipline this requires

Two things make the method work, and both are uncomfortable.

The first is that absence has to be recorded rather than resolved. When you look for evidence and do not find it, the honest entry is unknown, and here is why it stayed unknown. An unknown silently converted into an assumption is not a small bookkeeping lapse; it is the same silent-plausibility failure committed by the person doing the checking.

The second is that negative results have to be published as results. An audit that only records what it confirmed is not an audit, because you cannot tell whether it looked. The audit this comes from keeps a list of its own author’s prior conclusions that later measurement killed, four claims of absence that failed verification, and one apparent improvement written off as noise when it would not replicate. That list is the part that makes the rest worth reading.

What this is actually about

The reason this matters outside software is that the signature generalises without modification.

A report that is well-formed, on time, internally consistent, and built from a source that does not say what the report says it says. A process that produces a correctly formatted decision from an input nobody checked. A metric computed faithfully from a field that has been empty since a migration two years ago. Wrong input, well-formed output, nothing raised.

Fluency is not evidence. It is the absence of friction, and friction is what verification has always run on. When a system stops producing friction, the thing to check is not whether it is passing. It is whether it is looking.