Report · first published here

Do kept ideas come with a way to check them? One count, and what its first nonzero meant

A count the harness reads once per run asks how many of the ideas the judge kept both drew on three or more sources and carried a check a machine could use. Four runs in August 2026 read zero; the run that ended August 29 read 16 of 43, after the judge was allowed to write a check at the moment it kept an idea. Those 16 checks had not been run, and the source count reaches through inherited and grouped sources; as of October 2, 2026 no registered count covers kept ideas whose check was run and passed.

  • kept entries
  • checks
  • counting
  • the notebook

An idea is kept when the judge, Claude Opus 5 (Anthropic), the frontier model that reads batches of the small models' lines, chooses it and writes it up as a notebook entry. One count, read at the end of each run, asks how many of a run's kept entries both drew on at least three distinct sources and carried a check a machine could use; the methods paper calls it K3. A run is a period of operation under fixed settings; most settings are frozen for a run, a few (such as the filters' thresholds and the judge's model) can change inside one, and the harness tracks those changes. Its first five readings, from August 18 to 29, 2026, were 0 of 19, 0 of 41, 0 of 12, 0 of 45 and 16 of 43: four zeros and a nonzero, which is what this page's address refers to.

Why the count exists

The count was registered on August 19, 2026 as a cheap early signal for a stricter standard that had found nothing in seven runs and asks six things of an entry: at least three distinct sources; two or more steps over at least an hour; a form a machine can use; every step attributed to its maker; a pass-or-fail check that code ran; and no edit by a person or a session. That standard exists so that a planned comparison of the harness against cheaper ways of doing the same work has a unit to count. The count has no pass bar: two zeros in a row oblige the next session to register a new mechanism, and a nonzero reading permits no conclusion.

What was done between the zeros

The first two zeros invoked that rule. On August 22, 2026 a Claude Fable 5 session (Anthropic) registered and built the first mechanism: a quarter of the places for programs in each run went to kept entries already resting on three sources, and a program that passed on a conjecture lacking a check gave it one. When the next run also read zero, the session wrote down a fair test: the run of August 23 to 27 got triple the usual number of programs, and a second mechanism was owed if it read zero. It did. Almost none of the entries kept in the two runs before carried a check or rested on a line shaped like one, so the session of August 27 dropped a design that would extract checks from their text and registered one letting the judge write a runnable computation as it keeps an idea, omission being the honest default. The run of August 27 to 29 was the first under both, alongside other changes.

What the 16 are

In that run the judge wrote a check on all 43 entries it kept. The count accepts a written check only on an entry of conjecture type, and 26 of the 43 passed that clause. None of those checks had been run when the count was taken, so the 16 that also passed the source clause rest on written checks only. The run's form test for written checks (at least 40 characters and a comparison word) passed all 43, and the next run's audit found it cannot fail on checks hundreds of characters long. A fresh session of a Claude model (Anthropic) reproduced the count exactly from the raw rows.

The source clause is loose as well. The count finds sources by walking back through an entry's whole recorded lineage, including the grouping links a consolidation pass writes (a frontier-model pass that groups related entries under a theme), so an entry built from a theme reaches every member. A version of the stricter standard registered on August 27 counts only an entry's immediate parents, and in the same run none of its 37 candidates passed that clause. Without the grouping links the next run's count fell from 31 to 19; no reading of the 16 without them exists.

Since then

In the five runs through September 27, 2026, the count read 31 of 70 (19 without the grouping links), 30 of 72 (18), 29 of 53 (18), 60 of 106 (40) and 96 of 149 (64). As of October 2, 2026, no registered reading counts kept entries whose check was run and passed, and the stricter standard has not been met: on September 27, none of 111 candidates qualified.

What this report does not show

No mathematics is claimed from the count. It reads as a success only when its two caveats are left out (the counted checks had not been run, and its sources reach through inherited and grouped entries), which is why they stand beside the number.