Methods paper · dated August 29, 2026 · first published here
Self-Grading Without Self-Deception: Verdict Discipline for a Long-Running Mathematical Research Agent
This paper describes the rules under which Hypnos, a research harness whose loop has run since August 11, 2026, with pauses between runs, measures its own parts and records verdicts on them, and what those rules caught in its first 18 days. Claude Fable 5, an AI model made by Anthropic, wrote it on August 29, 2026. As of October 3, 2026, it has not been peer reviewed, and no human mathematician has read it.
- research methods
- preregistration
- measurement
- adversarial audit
- research agents
This is the methods paper of the Hypnos Math set, five papers written by AI models at the direction of David Ross: the other four are mathematics manuscripts, and this one is about how Hypnos, the research harness described above, graded its own parts from August 11 to 29, 2026, its first 18 days. Claude Fable 5, an AI model made by Anthropic, wrote it on August 29, 2026. The PDF has six pages, four of main text in a research workshop's format; it was never submitted to that workshop, prints no date, and has been published only here, since September 22, 2026. As of October 3, 2026, it has not been peer reviewed, and no human mathematician has read it.
Why a system that grades itself needs rules
A system grading its own work, the paper says, can fool itself with no visible error: the bar moves after the data arrive, negative results drop out of sight, and a broken instrument's silence reads as a rejection or a zero. Experimental science built preregistration against these failures, which the paper argues are sharper for a system that can rewrite its own instruments.
The parts it measures belong to the always-on loop, on mathematics the paper calls adjacent to the Riemann hypothesis: three small open models (Gemma 4 31B and Gemma 4 26B, made by Google, and Qwen3 32B) write one-line proposals; the judge, Claude Opus 5 (Anthropic), writes each one it keeps into the notebook as an entry of its own; and in an execution step separate calls to the judge's model write checks and then a program for some kept ideas, and code grades the result. The paper's unit is the run, a period of operation under fixed settings: most settings are frozen for a run, a few (the filters' thresholds, the weights that steer which entries the small models see, the judge's model) can change inside one, and the harness tracks those changes. Each measured part is on or off for a whole run, so each rate belongs to one configuration. Changes are made between runs, in working sessions that the owner, the person who runs Hypnos, starts by hand.
Six rules, and what each prevents
- Register a measurement before building it: a document filed before the code exists fixes the question, the pass bar, the least data needed to grade, exactly what the p-value tests and what each outcome sets in motion, so no bar is fitted to the data.
- Count "could not measure" as an outcome: a broken or dead instrument, or one short of data, yields COULD-NOT-MEASURE, which moves nothing, and no minimum is lowered to turn it into a verdict; every zero is checked for a dead instrument before it is believed.
- Reach a decisive verdict once, by a command that refuses a second run, so no result is rerun until noise gives a better-looking answer.
- Move a frozen number only forward, through a new registration that counts only later data while the old one stands, so no bar is edited to fit.
- Let kills run in both directions: a measured part moves at most one step per run along a ladder of four states (pre-registered, earning, probation, killed), so no run moves it twice, and a defective kill can be reversed with the reason recorded.
- Hold headline results for the end of the run: a ceiling registered in advance holds them for judging then, never by the session that produced them; a would-be headline, "refutes Riesz-basis stability", was held this way, and the paper draws no mathematical conclusion from it.
What the paper measured
Every number in the paper was copied from the project's files or database as they stood on August 29; none was recomputed. In 18 days the harness completed 12 runs: the small models wrote 221,787 proposals, the judge kept 313, and 87 execution jobs were dispatched. The judge reads only some proposals: in the paper's last run, August 27 to 29, it read 1,269 and kept 43, about 3 percent of those it read. From August 11 to August 29, 2026, about one proposal in 710 of all those written became an entry.
A standing adversary: 36 confirmations, five catches
The paper's seventh rule: at the end of every run since August 22, 2026, a fresh session of a Claude model (Anthropic), with no access to the working session's reasoning, recomputed every graded result from the raw data before any verdict was believed. The PDF's first table is its complete record to the cutoff, six audits: at three run ends it confirmed every graded verdict, 9, 13 and 14 of them, and it caught five faults.
- August 22: a defective kill. A module was killed after reaching probation on a COULD-NOT-MEASURE result, which may move nothing; the kill was reversed, and a graded miss killed it legitimately a run later.
- August 22: a p-value that ran the wrong way. A one-time verdict on storing the main small model's weights at 4 bits tested the model against its control, not against its own earlier rate, so a collapse to the control's level would have printed p = 1; it was re-registered before enough data existed to steer the change.
- August 25: two p-values for tests their registrations did not name, in the two one-time verdicts on the main small model (Gemma 4 31B over 26B, then 4-bit weights); notes were filed, the verdicts stand, and the naming rule was added.
- August 29: a group of jobs folded into another's pass rate in the run's report; no verdict had used it, and the code was fixed.
- August 29: a recount that could read entries filed after the run ended; every recount now cuts the data at the run's end.
In every case caught, the working session's own tests had passed, because they tested the wrong instrument; the paper takes the 36 agreements seriously because the same audit could disagree, and did.
The project's first finding was its own bug
On August 11, 2026, the day the harness was first built, one of its four numerical instruments made a mean of a squared magnitude negative, and the session building it blamed the formula it used, Theorem 2 of Feng's 2012 paper "Zeros of the Riemann zeta function on the critical line" (arXiv 1003.0059). On August 17, before any check ran, a Claude Fable 5 session wrote down five criteria that would kill that finding; of seven fresh sessions that tested it, three implemented the formula and found it correct as printed. The fault was a one-line bug in the project's own evaluator, repaired on August 18, before anything left the project. The note on this case tells it in full.
A kill audited into a better measurement
A module that added stray words to some of the small models' prompts was killed by its own gate, a bar on the share of proposals passing the harness's filters, set against a ceiling of 99.7 percent, which could show a broken pipeline but not what the module did. A post-mortem registered before it ran found that the module had moved what the small models wrote: the energy distance between the embeddings of batches with and without it gave p = 0.0025. The kill stood, and a new registration measuring that effect passed twice, with p = 0.0005, the smallest value its test can report.
What the end-to-end count shows
The paper's end-to-end example is the history of a count of kept entries that drew on at least three distinct earlier sources and carried a check a machine could use. The count has no pass bar, but two zeros in a row oblige the next working session to register a new mechanism. It read 0 of 19 and 0 of 41; then, with job slots reserved for kept entries with three sources, 0 of 12 and 0 of 45; then, once the judge was allowed to write a runnable check when it keeps an entry, 16 of 43. Every check counted was one the judge had written, not a program that ran, so the paper takes nothing from the 16: its receipt is the sequence, in which the loop measured its own miss, registered fixes on the record, set in advance the test that would oblige more, and the count moved. The next run's audit, on September 2, 2026, found that the form test those written checks had passed (at least 40 characters and a comparison word) cannot fail. The report on this count tells it run by run.
What has changed since the PDF
The PDF is the text of August 29, 2026. Four of its statements have since been corrected or overtaken:
- The "verified honest" zero. The PDF's example of an honest zero is a parser, looking for numerical bounds concerning the zeros of the Riemann zeta function, that filed none in its first live run; the PDF says the zero "was verified honest". That was corrected on September 2, 2026. The audit of the next run, whose one parse had been refused at the write, found that the check had covered the parser but not that write to the database, which was dead. A replay found no parse in the paper's last run, so the zero the PDF reports was a true count; the verdict stands, and the write was repaired. Why could-not-measure exists tells the case.
- The live diagram. Footnote 2 links a live observatory at an address that is not on this site; it now forwards to this site's live page on the loop, and the site's live figures give the current counts.
- The promise to publish the files. The footnote and the second table's caption promise to publish the registrations, run reports and other files the paper quotes. They are unpublished: they and the paper's reviews are in the project's private repository, and this site publishes reviewed explanations and selected receipts drawn from them.
- The cost figure. The dollar figure in the abstract and Section 6 is not a bill but the harness's own meter of frontier-model work priced at list rates; those models run on subscriptions, some calls were paid per token, and the project's files do not separate the two.
How the paper was made
Claude Fable 5 wrote the paper in one working session on August 29, 2026, at David Ross's request; the project's working sessions ran on it in August, and from September on its successor, Claude Fable 5.1, which wrote three of the four mathematics papers of the set. The PDF's author block, which predates the set's house format, prints two names, Claude Fable 5 (Anthropic) and David Ross, with a footnote saying the loop, its audits and the text are the machine's work, and David Ross directed the program and takes responsibility for its communication. This site cites it as it cites every paper of the set: Claude Fable 5 (Anthropic), at the direction of David Ross. He approved its publication here on September 22, 2026.
How it was reviewed
It had four reviews, all by AI models on August 29, 2026, each answered item by item that evening:
- A fresh instance of the writer's own model, Claude Fable 5, re-derived every number and statement; all 14 of its findings were fixed.
- GPT-5.6 (OpenAI), the other company's model and the earlier model of the owner's OpenAI agent, read it for errors and clarity; its 13 findings were checked and fixed.
- Three fresh instances of Claude Opus (Anthropic; the version is not recorded), a different model from the writer's company, read an anonymized copy with no knowledge of the project. They found the mechanisms described but not specified and every number testimony until the files behind it can be inspected, and called the adversary's sharing the judge's maker the sharpest objection to its independence.
- GPT-5.6 re-read the revised paper; its three small edits were applied, and that text is the PDF's.
What the paper does not claim
No mathematical discovery is claimed. The paper states three limits: every verdict is for one loop and one area of mathematics; the judge is a frontier model whose instruction following its maker tuned, so the controls sit outside it (the judge writes checks, and code decides whether a machine can check them); and the adversary comes from the judge's own maker and was installed partway through, so the runs before August 22 were never audited this way. The paper's thesis, that this discipline rather than the models' capability is the reliability layer such a loop is missing, is argued from one loop's history, not tested; one reviewer put it as "n=1 loop, no ablation".
If you read this paper
The owner would like to know what you make of it. A few words are enough: right, wrong, known, minor, worth a look, or not your area. In his words: "If it is right I want it in the record." If it is wrong, the page is corrected the same day and you are told; if it is already known, the page credits the earlier work.
Write to david@hypnosmath.org. He reads and answers replies himself; an error you find is fixed in the published files and on this site the same day, and you are told; nothing you send is given to any AI model without your permission.
The file
A SHA-256 fingerprint identifies a file exactly: a copy with the fingerprint below is, byte for byte, the PDF this page describes (69,911 bytes), the only file published for this paper.
The PDF has 6 pages and SHA-256:
07adc64f3bf73ab29a42214a8bd33e099fb8eecdee0eb4c9beb8dadfb5b15be5