Report · first published here

What the record showed by September 22, 2026, and what it has shown since

An assessment made on September 22, 2026 found that the harness's record held real but narrow tooling, reproductions of known mathematics and corrections to its own measuring code, and did not establish a new mathematical discovery or useful accumulation by the loop. The four mathematics papers came later, were written by frontier models, and are outside it.

  • assessment
  • the live loop
  • what the record shows
  • dated verdict

On September 22, 2026, a session of OpenAI's Codex agent (the command-line program through which the owner, the person who runs Hypnos, uses OpenAI's models) read the harness's private record without changing it or running new experiments. What it reviewed was the harness's record as it stood on September 22, 2026, including the work of the live loop, the part of Hypnos that runs on its own (small open models write candidate ideas, a frontier model keeps a few, and programs test some of those); the four mathematics papers in the site's library came later, were written by frontier models, and are outside this assessment. The evidence reviewed here does not establish a new mathematical discovery or useful autonomous accumulation by the live loop.

The record did show real but narrow tooling for reusing a checked result, reproductions of known mathematics, and corrections to the project's own measuring code. Useful accumulation means later work demonstrably using earlier checked results: removing the earlier result breaks the later one. Some reuse existed (a review GPT-6 Astra wrote on September 12, 2026 in a Codex session counted 28 later entries built on 19 program results), so the open question was usefulness.

Three things the record did show

A tool that saves a checked result for reuse, September 13, 2026. The same Codex session built a tool that saves the result of one exact computation with a certificate; in a test registered in advance, two later checks each verified and used it, damaged or missing copies were refused, and reuse was slower than recomputing. The test made no model calls, and nothing in the running loop calls the tool.

A published formula that was right, August 17 to 18, 2026. When a project program made a published formula look wrong, checks run under criteria that a Claude Fable 5 session (Anthropic) wrote down beforehand found the formula right and the fault a one-line bug in the project's own code. Nothing went outside the project.

A benchmark that lost seven of its rows, September 21 to 22, 2026. GPT-6 Astra (OpenAI), auditing the harness, found that code reading a fixed benchmark of Claude Opus 5's past decisions dropped seven rows; a Claude Fable 5.1 session (Anthropic) verified and fixed it the next day. The benchmark's verdict of August 2026, which restored a cheap score's role in ranking what Claude Opus 5 reads, stood, since both counts clear the same bar.

What the record did not support

A kept entry means that the judge, Claude Opus 5 (Anthropic), the frontier model that reads batches of the small models' lines and writes each kept one up as an entry, found it worth keeping, not that it is true or new. A program the harness ran shows only what its own check tested.

A count of kept entries that both rest on three or more sources and carry a check a machine could use read zero four times in August 2026, then 16 of 43 on August 29, and those 16 carried checks the judge had written, not checks that had run.

A review of September 1 to 2, 2026 by Claude Fable 5.1 sessions ruled two flagged computations valid but not firsts: one repeated an earlier program's value, and the other rested on an assumed conjecture.

Two other models read what the judge read, with no power over any item: from September 14 to 19, GPT-6 Astra would have kept 2 of the 106 ideas the judge kept, and Gemma 4 31B, one of the three small models, 7 of 105. Astra's overall agreement of 0.965 was below the 0.967 of a reader rejecting everything. Shown 118 of its own earlier rejections again, the judge kept 9. No comparison here has an independent label of which ideas are good.

What has happened since September 22

From September 28 to October 1, 2026, in sessions the owner started, Claude Fable 5.1 wrote three mathematics papers and GPT-6 Astra one, each proving theorems on a problem its writer chose. Their origins differ, and the library gives each in full: only the unfolded-zeros paper's chain begins with a small model's line, a recall of Kadec's 1/4 theorem kept by the judge as a question on August 12, 2026; the antiunitary paper's main theorem is its writing session's reformulation, motivated by five notebook records; the Gil note answers two questions Gil posed in his own paper, found in a review of the antiunitary paper, with none of its mathematics from the harness; the dividing-plane note came from a line of work the owner ran with frontier-model sessions on the OpenAI forced Navier-Stokes blow-up manuscript. As of October 3, 2026, no human mathematician has read any of them.

Each paper was reviewed only by AI models, each round finding items earlier rounds had missed; the record does not show the error rate falling to zero. Each presents its main result as its own contribution after a bounded literature search, and says that the search does not establish priority.

On September 27, 2026 the judge was tested on planted errors, a test the assessment had listed as not yet run: 24 kept entries, each altered by one deliberate error, were sent to it among their originals. It kept 11 of the altered ones under its rules of the time and 1 under new rules that make it mark each item on five points first, rules that also rejected 13 of the 24 originals against 10. The test had no pass bar, and the new rules are now in use.

As of October 2, 2026, the harness has no registered measurement of kept entries whose written check later ran and passed, and on September 27 none of 111 candidates met its stricter standard, whose six conditions include three direct sources, a check that code ran and passed, and no edit by hand.

What would change the answer

The record describes the test that would settle the open question, whether checked work can become useful progress on later tasks without the system merely being rewarded for its own judgments. In it, a checked result from one task removes a named obstacle in another that reads it, across independent families of tasks, with a comparison against recomputing registered in advance and every attempt counted.

What this report does not show

None of these figures can be recomputed from this site; the harness's notebook, logs and database are private. Whether the layer of small models saves frontier-model effort is untested. The assessment is one session's reading of the record.