Note · first published here

What Hypnos does with the ideas it turns down

From August 11 to September 27, 2026 the small models wrote 485,873 candidates; Claude Opus 5, the frontier model that judges them, read about one in twenty and kept about 3 in 100 of those it read; the rest were not deleted. A rejection is stored with its reason and the premises it leaned on, and is put back in the queue when a later kept entry opposes one of them; a failed check is scoped to the instance it tested; a check that could not decide says so instead of saying no; a result that would be a first is held for a later ruling. The run reports show the machinery in use and one rejection reopened by it.

  • rejections
  • scope
  • could-not-measure

Most lines are turned down, and none is deleted

From August 11 to September 27, 2026, the small models wrote 485,873 candidates. The judge, Claude Opus 5 (Anthropic), the frontier model that reads batches of the small models' lines and writes each kept one up as an entry, read 22,464 of them, about one in twenty; the rest waited in a queue or were set aside when a queue was full, never deleted. It kept 762 of those it read, about 3 in 100. So about one candidate in 640 became a notebook entry. Lines set aside keep their text and scores, and every batch the judge reads includes a few drawn from them at random.

Why several kinds of no

The harness's design reads a rejection narrowly: this instance failed against this target, in this context, as measured by this metric. It names four ways that reading gets inflated: from one instance to a whole type, from one target to all targets, from today's inputs to forever, and from an invalid test to any verdict. The rules against them were written on August 11, 2026, in answer to a question from the owner, the person who runs Hypnos, about a kind of idea being ruled out everywhere after failing in one place. The list of premises behind each rejection followed on August 14, from his concern that the judge might turn down an idea that fails against the notebook as it stands but would hold once a missing piece was in place.

The kinds, and what each leads to

A rejection by the judge is stored with its reason and the premises it leaned on; when a later kept entry opposes one of them, code returns every rejection that leaned on it to the queue with that entry attached, and no model re-judges the archive. From August 27 to September 27, 2026, between 1,120 and 4,049 rejections in each run (a run is a period of operation with most settings frozen) recorded their premises, and one rejection was reopened this way, in the run of August 29 to September 1. The run reports do not say what came of it.

Pass and fail, a check's two decided outcomes, are computed by code from constants frozen in advance, never by a model. A failed computation is stored with what was tested, the target, the premises in force and a scope that defaults to this instance on this target: "this computation refuted this idea under these premises" is not "the idea is dead". Under the design, rejections go stale when a premise changes, and a near miss can wait with written conditions that wake it. Could-not-measure is a third outcome, not a no.

Quarantine holds a result that would be a first (it beats a certified value, contradicts a cited bound or exceeds a ceiling set in advance): it files nothing as a finding and is judged only at the end of a run, never by the session that produced it. The design calls a quarantined result the best outcome a program can produce. In one case a ceiling held back a would-be refutation of Riesz-basis stability; the ruling upheld the ceiling, and no mathematical statement is made from it.

A verdict is also locked to what it tested: the two verdicts that moved one small-model slot to a larger model and then to its 4-bit version say nothing about model size or quantization in general. Each measured part of the harness has a standing that one run can move at most once, and a kill, which switches a part off, can be undone: one was reversed when an audit found that the part's earlier step down had come from a could-not-measure, which may move nothing; the same part was later killed by a graded miss.

What this note shows and does not show

The run reports show premises recorded and one rejection reopened. Nothing here shows that any rejected idea later turned out to be valuable, or that the judge's rejections are correct: a blind re-read of kept and rejected texts found no significant separation at its resolving power.