Note · first published here

The formula was right: the project's first finding, and the hour that ended it

On August 11, 2026, the day the harness was first built, a session computed a negative value for a quantity that cannot be negative and concluded that a printed formula behind a published bound on the zeros of the zeta function was wrong. On August 17 a session wrote down, before any check ran, what would abandon that finding; seven separate sessions, each started fresh, ran the checks; within an hour three of five criteria went against the finding, and the cause was a one-line bug in the project's own program. Nothing was sent outside the project.

  • preregistration
  • reproducibility
  • verification

What happened

The project's first finding, on August 11, 2026, was that a printed formula behind a published bound on the zeros of the zeta function was wrong. On August 17 a Claude Fable 5 session (Anthropic) tested it with checks written down before they ran, and within an hour found the formula correct as printed and the error in one line of the project's own program, before anything left the project.

The formula, and why a negative value was impossible

The bound is on κ, the proportion of zeros of the Riemann zeta function on the critical line. Pratt, Robles, Zaharescu and Zeindler (arXiv 1802.10521, 2020; PRZZ below) state that κ > 0.417293 and that the proportion κ* of simple zeros on the line is at least 0.407511. The computation behind such a bound evaluates a mean value c, which the papers present as an average of a squared magnitude, so c cannot be negative; the bound, κ ≥ 1 − log(c)/R for a parameter R, needs c > 0. The project took its formula for c from Theorem 2 of Feng's "Zeros of the Riemann zeta function on the critical line" (arXiv 1003.0059, 2012); PRZZ write that their main-term code matched Feng's.

The harness's first build, on the afternoon of August 11, before the loop first ran that evening, made four numerical tools and required each to reproduce a published result before research could start; the fourth was to reproduce PRZZ's bound from their printed polynomials. Its program for the formula, written by a Claude Fable 5 session, found c negative at both parameter sets in use, PRZZ's and the one in Feng's Corollary 1, and the session concluded that the printed formula could not have produced the published numbers, wrote a mathematical explanation for that, and wrote a test that expected the negative value. The same afternoon a Claude Fable 5 session reproduced PRZZ's κ and κ* to all nine printed digits by a separate route, from the PRZZ paper's sections 6 and 7.

Checks written down before they ran

On August 17 a research digest by a Claude Fable 5 review session called the finding unreported and worth publishing, and a session was queued to write it up. That session turned the task around: the hypothesis under test became "this finding is publishable", and its job was to refute it. That evening, before any check ran, it recorded five kill criteria (the project's term for tests that state in advance which result abandons a finding), each with a fixed bar, under the rule that no bar moves afterward.

Seven separate sessions, each started fresh, ran the checks: two built the formula from the papers alone, by different methods; one argued for the paper, trying every reading of the printed text an expert might use; one applied the same kind of computation to a chain of formulas known to be correct; two searched the literature; one checked from the papers' definitions that the mean value cannot be negative. None read the project's code or another session's output, and the two blind builders were given no expected sign or value. The criteria document grounds this separation in a principle, that bias correction is structural, never introspective, and names no model for those sessions.

Three of the sessions implemented the printed formula, the two blind builders and the advocate, and all three agree digit for digit that c is positive: 2.137450131107766 at PRZZ's parameters and 1.352926843466767 at Feng's. The advocate found that only the literal reading reproduces anything published, and that reading reproduced four of Feng's printed values, and PRZZ's κ* to all nine printed digits. One leg of one criterion had no published value to compare with and was recorded as could-not-measure. The positivity premise survived: the papers' definitions force c to be positive, so the impossibility argument was sound and its input wrong. The verdict came 43 minutes after the criteria: three of the five went against the finding, and the formula is correct as printed.

The cause

A table of coefficients in the project's program for the formula was filled only for pairs of indices whose sum is at most 5, where the formula needs every pair with each index at most 5: a one-line bug. It has no effect when the polynomial in the formula has degree 2 or less and gives wrong values from degree 3 up; both parameter sets use polynomials of degree 5. The August 11 session's two cross-checks had agreed on the negative value because both used the same table, and the verification's report calls the August 11 explanation "a story fitted to an index bound".

The session that ran the verification repaired the program in the early hours of August 18, 2026, at the direction of the owner, the person who runs Hypnos: it fixed the line, replaced the test that expected the negative value with tests of the correct positive values, and corrected the false explanation. The program now matches the independent implementations to machine precision, and the tool passes at κ = 0.417293962029. The notes of August 11 and 17 were left as written and corrected by later dated entries.

What it shows and does not show

Checks written down before they ran, and built to share nothing with the original program, overturned the project's own finding within an hour; checks that share an intermediate calculation are not independent at that point. The methods paper tells the episode, with its three implementations, as the opening self-refutation: the machinery for overturning the project's own results ran before any result worth keeping existed. A working session did the verification and the repair, not the always-on loop of small models, and they show no new mathematics: the repaired program reproduces known values. The criteria, the report, the sessions' reports and the log of the repair are in the project's private repository; the public account is this note, the methods paper's section and the worked case.