Note · first published here

Why a check that finds nothing says 'could not measure', not 'no'

A check that could not measure (too little data, no opportunity, a dead part, broken tooling) records that it could not, moves nothing, and is never turned into a verdict by lowering its floor. The rule came from the owner on August 11, 2026, before anything was built. In the case told here, a screen read zero in two runs: the first zero reading was true, the second was a write the database had refused, found by a later audit, and the outcome was the same either way.

  • could-not-measure
  • preregistration
  • audit

When a check finds nothing

Hypnos does not read a check that finds nothing as a no. Its checks speak three words: pass, fail and could-not-measure. The third is for a check that could not do its job: too little data to clear its floor (the least it needs to grade), no chance to observe, a part it depends on down, or broken tooling. Could-not-measure is a state, not a verdict: it changes nothing and is never turned into a verdict by lowering the floor. A broken instrument must never emit a rejection, and since a dead instrument also reads zero, a zero reading is checked for a dead instrument before it is believed.

Where the rule came from

On August 11, 2026, before anything was built, the owner, the person who runs Hypnos, required, in his words, "testing the build before starting the real work, so it doesn't rule things out based on false indications"; the first build, that day, used the three words from the start.

Other checks end the same way: one leg of one criterion in the verification of August 17, 2026, told in the note on the formula that was right, had no published value to read against; and a program testing a kept idea that misses its known value is could-not-measure about the program, never about the idea.

The case: a screen that read zero

The case is a screen, a program with no language model in it, that reads every line the small models write for a stated numeric bound on one of five quantities about the zeros of the Riemann zeta function and compares it with the best known value. It removes, ranks or blocks nothing. Its one graded test comes at the end of each run, a run being a period of operation under fixed settings (most settings are frozen for a run, a few can change inside one, and the harness tracks those changes). Up to 60 of the screen's matches go, as text alone, to a second model, which this note's sources do not name, to say what bound each states; at least 90 percent must match the screen's reading. Below 20 matches in a run the test records could-not-measure, and that floor is never lowered.

A Claude Fable 5 session (Anthropic) registered the screen on August 25, 2026, writing its test, bar and floor down before any of its code existed, and wrote that a first run below the floor would be the design working, not failing. During the first run, August 27 to 29, a Claude Fable 5 session checked for a dead instrument: the screen was switched on and in use, had logged no error, and its parser answered test sentences and refused a deliberately loose one. It called the zero reading honest. That run and the next, August 29 to September 1, ended could-not-measure with no matches.

The check had exercised the parser but not the step that saves a match, and that step was broken: the small-model loop's database connection refuses writes to any table not on an allowed list, and the screen's table had never been added. An audit at the end of the second run found it: a fresh session of a Claude model (Anthropic), with no access to the reasoning behind the run's results, re-derived them from the raw data. Replayed over that run's 23,157 lines, the parser found exactly one match, at the moment of the screen's only error event, which the run report had not shown; over the first run's 20,020 lines it found none. So the first zero reading was a true count, and the second was a write the database had refused. Either way the count was below the floor of 20, so the outcome stood; only its stated cause was wrong. On September 2, 2026 a Claude Fable 5.1 session (Anthropic) added the table to the list, added a test that writes through that connection, and made the run report show every kind of error event.

The registration treats the yield, one match in 43,177 lines against the tens per run it expected, as information about what the small models write, and keeps the floor. The next four runs, through September 27, 2026, filed 2, 1, 3 and 3 matches, so by then the test had never graded.

What the case shows and does not show

It shows a check that found nothing recording could-not-measure, as its registration said it might, and a dead-instrument check that was itself incomplete, caught by a later independent audit with the outcome unchanged. It shows nothing about mathematics, nor whether the screen reads stated bounds correctly. It is one case; it does not show that every zero reading has been checked end to end.