Worked case: An engineering test of reuse

A checked calculation saved once and verified by two later checks

On September 13, 2026, a GPT-6 Astra (OpenAI) session built a tool for Hypnos that saves one exactly checked calculation and makes every later use verify it before relying on it. In a test written down in advance, two later checks each verified one saved result, and every damaged copy was refused; the test made no model calls and ran in under two seconds. Nothing in the harness's running loop calls the tool.

Hypnos is a research harness with two parts. In the loop, which has run since August 11, 2026, with pauses between runs, three small open AI models on the graphics cards of the owner, the person who runs Hypnos, write one-line ideas about entries in a notebook of mathematics; Claude Opus 5 (Anthropic), the judge, reads samples and writes the few worth a closer look into the notebook as entries. In sessions the owner starts by hand, Claude Fable 5.1 (Anthropic) and GPT-6 Astra (OpenAI) read chains of entries and do the deriving and the checking; whether the small models save them any effort is untested. How it works, in full

This case is an engineering test that a session designed, built and ran, not a step of the always-on loop. The report this case links to is its published account; the file the comparison wrote is private.
This case runs from September 1 to 14, 2026.

Six steps and the five links between them

Each step is one piece of the work; each link between steps says how one step's output became the next one's input. Every step names who did the work, by model and maker, as far as the published sources record it; they do not say which small model wrote a line. Each panel links to the published page that reports the case.

Six steps and five links, in order

Step 1 · September 12, 2026

A review finds no route from a checked result to later work

On September 12, 2026, the owner asked OpenAI's Codex agent, the command-line program through which GPT-6 Astra (OpenAI) runs, to read the OpenAI forced Navier-Stokes blow-up manuscript, released on September 8, and say what Hypnos could learn from it. The review it wrote is about methods, not mathematics: what it takes from the manuscript is the practice of breaking work into small exact results that later work can check and reuse. Its central finding about Hypnos: the harness had machinery for variety, memory and checking, but no reliable route from a checked partial result to continued work on it, because a program of the execution step receives an idea's description and nothing that loads the programs and results behind it. Some reuse happened anyway: 19 program results had 28 later entries derived from them, and 9 later programs ran on those.

Started from
The OpenAI manuscript, Hypnos's code and log, and the owner's question.
Done by
GPT-6 Astra (OpenAI).
Produced
A written review, and one piece to test its advice on: a small exact lemma from a notebook entry.

From step 1 to step 2 · Narrowed

From a 166-page manuscript to a lemma that fits on a page

To try its advice, the review needed one small exact result with a later use. It took a notebook entry filed by a session ten days earlier, whose own statement was broad and had no program run on it yet, and cut out the one part with an exact answer: the eigenvalue counts when every displacement is +1 or −1.

Step 2 · September 1 to 12, 2026

A small exact lemma, with its proof

Take distinct integer positions j, give each the value u_j = +1 or −1, and form the symmetric matrix M with entries 4π²(−1)^(j−k)(u_j − u_k)/(j − k) off the diagonal and 0 on it. The lemma: with p positions at +1 and q at −1, M has min(p, q) negative eigenvalues, |p − q| zero eigenvalues and min(p, q) positive ones. The proof: changing the signs of rows and columns removes the alternating factor, entries between positions with equal values vanish, and ordering the +1 positions first leaves a block form whose off-diagonal block C, with entries 2/(x_a − y_b), is a Cauchy matrix of full rank min(p, q); each nonzero singular value of C gives one positive and one negative eigenvalue. Five positions at +1 and two at −1, for example, give the counts (2, 3, 2). The ingredients are standard; the review calls isolating the lemma an application of method. M is the derivative, at zero displacement, of a quadratic form an earlier program of the harness had computed: a Claude Fable 5.1 (Anthropic) session of September 1 to 2 derived that formula and filed it into the notebook, which sessions can write to and the small models cannot.

Started from
The derivative of an earlier program's quadratic form, filed into the notebook on September 2.
Done by
The derivation: Claude Fable 5.1 (Anthropic). Isolating the lemma: GPT-6 Astra (OpenAI), in the review.
Produced
One exact lemma with a proof, and a predicted count for every assignment of +1 and −1.

From step 2 to step 3 · Derived

The proof's shape is what the tool checks

The proof reduces the counts to the rank of one Cauchy block, so a saved result needs to carry only what certifies that rank: the block as exact fractions and a nonzero determinant. A result that passes the checks is exact for its case; the general statement rests on the proof, not on the saved files.

Step 3 · September 13, 2026

A tool that saves a checked result and makes every user re-check it

The tool has two programs. The first takes the positions and values, writes the block C as exact fractions, and computes the determinant of its leading square block by the Cauchy determinant formula, reduced modulo a large prime. The second, before using anything, re-checks every entry against the definition, recomputes the same determinant by elimination, a different calculation, requires it to be nonzero and equal to the first program's, and only then reads off the three counts. Each saved result is stored under the SHA-256 digest of its bytes; a user checks the digest, which says it has the intended file, and the mathematics, which says the file is right. Citing a saved result earns no credit for reusing it, and the tool calls no model.

Started from
The lemma and its proof, which the second program's checks follow.
Done by
GPT-6 Astra (OpenAI), at the owner's request.
Produced
Two programs and a store of saved results, each keyed by its digest.

From step 3 to step 4 · Run

Two ways of working, the same seven cases, no model involved

The comparison ran the two ways of working on the same cases: rebuild the result for every question, or build it once and verify it at every use. Its design, written first, said that checking might cost more than recomputing and registered no expectation about speed. The run made no model calls, and its output file marks the work as the session's, not the unattended loop's.

Step 4 · September 13, 2026

One comparison, written down before it ran

The design was written down before the run and amended once, still before it, to correct two of its readouts. Seven cases: five valid assignments of +1 and −1 on 5 to 129 positions, and two invalid controls, duplicate positions and a value other than +1 or −1. Each case had two later questions: the three counts, and the number of zero eigenvalues alone, which the first contains. One way of working rebuilt the result for every question; the other built it once and handed it to both. The counts came out as the lemma says: (3, 0, 3), (2, 3, 2), (2, 1, 2), (0, 5, 0) and (63, 3, 63). Each way completed 10 of its 14 questions; the 4 on the invalid controls were refused both times, as designed, and are counted as not solved rather than dropped. Builds: 14 against 7. Time, builds included: 1.500 seconds against 1.812, so reuse was slower, because checking dominated so small a workload.

Started from
Seven fixed cases with their questions, written down before the run.
Done by
GPT-6 Astra (OpenAI), the same session; the run itself made no model calls.
Produced
A file of every count, time and refusal, kept in the project's private files.

From step 4 to step 5 · Verified

What the ten refusals show

Ten controls removed or corrupted the saved result, each valid case once each way, and all ten users refused to complete: each depended on the saved result rather than citing it. A deliberately wrong result was submitted first and failed; its failure was kept, and the corrected one completed the same question on the second attempt. All 10 completions that reused the saved result recorded a verified read of it.

Step 5 · September 13 to 14, 2026

Two later checks verified one saved result; damaged copies were refused

Within this fixed test, two later checks each read and verified the same saved result and completed, and every damaged or missing copy was refused. That is all it shows: no speedup, and nothing about any model's ability, since no model was called; the two questions are closely related ones within one narrow family of tools; and a saved result is exact evidence for the cases run, not a formal proof of the lemma. On September 14 a Claude Fable 5.1 (Anthropic) session reviewed the tool, and it became part of Hypnos's code that day.

Started from
The comparison's output file and the design that preceded it.
Done by
The review of September 14: Claude Fable 5.1 (Anthropic).
Produced
A tool in Hypnos's code, and a report of what the one test showed.

From step 5 to step 6 · Unconnected

A tool the loop does not use

The tool has sat beside the loop since September 14. The comparison in which a model-driven program would use checked work was not registered: its cost was referred to the owner, and through October 2, 2026 the harness's log holds no registration or result of it.

Step 6 · As recorded through October 2, 2026

Nothing in the running loop calls it

As of October 2, 2026, the tool runs only when a person types its command; no part of the always-on loop calls it. The test that would matter, a model-driven task using checked work to remove another task's real obstacle, measured on independent families of tasks with budgets, checks and stopping rules fixed in advance, has not been run. The lemma covers only the values +1 and −1 and the derivative at zero displacement; at a finite displacement, higher-order terms can split the zero directions.

Started from
Hypnos's code and log as of October 2, 2026.
Done by
Read from the code and the log; the loop played no part.
Produced
One open question: whether checked work removes another task's obstacle at a useful total cost.

What this case shows

Within one fixed test, one saved result was verified by two later checks and every damaged copy was refused; whether checked work can clear another task's real obstacle has not been tested.

Next

Other views of the same work