Worked case: A wrong finding, tested and withdrawn
Six days of a wrong finding, and the hour that overturned it
On August 11, 2026, a Claude Fable 5 (Anthropic) session building one of Hypnos's four numerical tools concluded that a formula printed in a published paper on the zeros of the Riemann zeta function was wrong: evaluated, it gave a negative value for an average of squares. The finding stood for six days. On August 17, five tests written down before any ran, carried out by seven fresh sessions that shared nothing with the tool, showed within the hour that the formula was right and the error was one line of Hypnos's own code. Sessions caught it, not the loop of small models, and nothing had left the project.
Hypnos is a research harness with two parts. In the loop, which has run since August 11, 2026, with pauses between runs, three small open AI models on the graphics cards of the owner, the person who runs Hypnos, write one-line ideas about entries in a notebook of mathematics; Claude Opus 5 (Anthropic), the judge, reads samples and writes the few worth a closer look into the notebook as entries. In sessions the owner starts by hand, Claude Fable 5.1 (Anthropic) and GPT-6 Astra (OpenAI) read chains of entries and do the deriving and the checking; whether the small models save them any effort is untested. How it works, in full
The written tests, the verdict and the checks' programs are in the project's private files and are not published. The note this case links to is the published account; the methods paper also describes the episode. This case runs from August 11 to 18, 2026.
Six steps and the five links between them
Each step is one piece of the work; each link between steps says how one step's output became the next one's input. Every step names who did the work, by model and maker, as far as the published sources record it; they do not say which small model wrote a line. Each panel links to the published page that reports the case.
Six steps and five links, in order
Step 1 · August 11, 2026
A new tool reports an impossible negative value
Before any research ran, Hypnos's first build made four computational instruments, each checked against published values on August 11, 2026. The fourth was to reproduce, from the polynomials printed in their paper, the lower bound of Pratt, Robles, Zaharescu and Zeindler (2020, arXiv 1802.10521): more than 41.7 percent of the zeros of the Riemann zeta function lie on the critical line. A result announced in August 2026 gives a higher bound, 67.25 percent, for zeros that are simple and on the line; the 2020 bound was this tool's target. Behind the bound is a mean value c, an average of a squared magnitude, so c cannot be negative. For c the tool used the formula printed as Theorem 2 of a 2012 paper by Feng (arXiv 1003.0059), which the 2020 paper says its own main-term code matched. The session recorded c = −7.04 at the 2020 paper's parameters and −92.8 at Feng's, concluded that the printed formula could not have produced the published numbers, wrote a mathematical explanation, and fixed the negative result in a test.
Started from
Feng's printed formula, and the polynomials and parameters printed in the two papers.
Done by
A Claude Fable 5 (Anthropic) session, building Hypnos's first version.
Produced
A recorded finding that the printed formula was wrong, with an explanation and a test.
Two checks agreed with the finding; neither could catch the bug
A negative average of squares admits two readings: the printed formula is wrong, or the program evaluating it is. The session took the first, and two checks seemed to agree: two tools it believed independent gave the same negative value, and the printed formula matched a refereed closed form of Conrey's to 30 digits. Both checks, it turned out later, were blind to the bug. For six days the project's log carried the finding with its explanation.
The tool meets its target, and a summary calls the defect unreported
The same afternoon a Claude Fable 5 session reproduced the 2020 paper's two published values to nine digits by a separate route through that paper; those two numbers were the tool's target, so it passed, and research could start. The negative value stayed in the project's log as a finding about Feng's formula. On the evening of August 17 a Claude Fable 5 session summarizing the project's research called it unclaimed territory, a defect nobody had reported, and a session was set up to write a short note on it.
Started from
The August 11 finding and its explanation, as the project's log held them.
Done by
Claude Fable 5 (Anthropic) sessions: one for the reproduction, one for the summary.
Produced
A finding ready to be written up for readers outside the project.
Asked to write it up, the session set out to refute it
The session given the note did not write it. It made the finding the thing under test: the statement was that the finding could be published, the session's job was to refute it, and a clean refutation would count as a fully successful outcome.
The session wrote five tests, each stating in advance which result would abandon the finding, with fixed bars and the exact instructions each check would receive, and fixed them in the project's log before any check ran: no test would be regraded after its first reading, and no bar would move. The design followed a principle the owner had set for the project: bias correction is structural, never introspective.
Started from
The finding, the two papers, and the principle the owner set.
Done by
The same Claude Fable 5 (Anthropic) session, before any check ran.
Produced
Five tests with fixed bars, and a plan for seven checks kept apart.
Following the owner's principle, independence was built in rather than asserted. Each new check got the two published papers and nothing from Hypnos, and saw no other check's output; the two that would build the formula were given no expected sign or value and were not told that a prior finding existed.
Seven fresh sessions, kept apart, check the formula
Seven fresh sessions ran the checks. Two built the formula from the papers alone, one with exact series arithmetic and one by numerical differentiation through Cauchy contour integrals. One argued for the paper, trying every reading of the printed text an expert might use. One applied the same method to formulas known to be correct, as a control. Two searched the published literature. One checked, from the papers' own definitions, that the mean value cannot be negative. The numbers were compared only after all seven had returned.
Started from
The two published papers and nothing from Hypnos.
Done by
Seven fresh sessions run by the Claude Fable 5 (Anthropic) session; the sources do not say which model they ran on.
Side by side, the results point at Hypnos's own code
Three of the seven, the two blind builds among them, had implemented the printed formula, and all three agreed digit for digit on a positive value. The session that argued for the paper found that only the literal reading reproduced the published digits, and the premise check held: the impossibility argument was sound, and its input was wrong. So the error had to be in Hypnos's own evaluator.
The formula was right; the bug was one line of Hypnos's code
Three of the five tests fired; one would have been enough. The printed formula gives c = 2.137450131107766 at the 2020 paper's parameters and 1.352926843466767 at Feng's, and the verdict was recorded 43 minutes after the tests were fixed. The cause: a table of series coefficients was filled only for index pairs whose sum was at most 5, where the formula needs every pair with each index at most 5. For polynomials of degree 2 or less the broken table gives the right answer, which is why the closed-form check had passed, and both earlier tools had used that one table. Early on August 18 the same session fixed the line, replaced the test that held the negative result with tests holding the correct values, and corrected the false explanation. The report's lesson: "Decorrelation that shares one ingredient is one tool."
Started from
The seven results, read against the tests written before them.
Done by
The Claude Fable 5 (Anthropic) session; the repair at the owner's direction.
Produced
A recorded verdict, a one-line repair, and a tool that reproduces the published values.
The loop of small models and the judge played no part. The finding came from a tool the first build made before research began, and sessions the owner started found it, tested it and repaired it. What caught it was a method: tests written down before the check, and checks kept apart.
It produced no mathematics: the repaired tool reproduces values already published, and no later result follows from the repair. Nothing was sent anywhere: no note, no contact with the authors, no submission. The tests cover this one finding; a defect reported later in the same formulas would need its own tests, written in advance. The methods paper reports the episode as the project's opening self-refutation: the machinery for refuting its own results ran before any result worth keeping existed.
Started from
The project's private log of the episode.
Done by
Claude Fable 5 (Anthropic) sessions throughout.
Produced
One corrected tool, and an account of a wrong finding that stood six days and fell within an hour of being tested.
Inside the project the wrong finding stood for six days; once tested, it fell in under an hour; nothing had left the project, and the tests were written down before they ran.