Report · first published here
The catches: when the harness's own audits found it wrong
At the end of each run from August 22 to September 22, 2026, a fresh session of a Claude model re-derived every graded number from the raw rows before any verdict was believed, and GPT-6 Astra did one such audit too. This page lists what they caught: a kill made against the harness's own rules, three significance figures bound to the wrong question, a group folded into another's pass rate, and faults in the code that reads and renders the record. No verdict was re-graded; each fault was fixed forward.
- audits
- self-checking
- measurement
- trust
At the end of each run, before any verdict was believed, a fresh session, with no access to the grading session's reasoning, re-derived every graded number from the raw rows and the frozen registrations with its own queries, marking each row reproduced, differing or not re-derivable. A run is a period of operation under fixed settings; most settings are frozen for a run, a few (such as the filters' thresholds and the reviewing model) can change inside one, and the harness tracks those changes. A registration is a design, written before the code it governs, that names what will be counted and the bar a result must clear. A one-time verdict is taken by a command that refuses to run twice. Fresh sessions of a Claude model (Anthropic) ran the audits at the end of every run from August 22 to September 22, 2026, and GPT-6 Astra (OpenAI) audited the run of September 14 to 19 a second time; the audit of the run that ended September 27 is written and has not run.
Why the record keeps this list
In every case caught through August 29, 2026, the grading session's own tests were passing, because they tested the wrong thing. An audit's agreements mean something only because the same procedure also produced real disagreements. From the audit of September 5 on, every catch was in code that reads or renders the record, none in a graded verdict, and no verdict was re-graded.
Six cases from August 2026
1. A kill that broke its own rule (August 22, 2026). Each measured part of the harness has a standing (registered, earning, on probation or killed). One part was killed after a graded miss, though its earlier step down had come from a could-not-measure outcome, which may not move a standing; the grading code had copied that step from other parts' registrations, and the old tests asserted it. The kill was reversed by a registered way of undoing one, the code and tests were fixed, and the part was killed legitimately at the next run's end.
2. A significance figure bound to the wrong question (August 22, 2026). A one-time verdict asked whether compressing one small model to 4 bits had cost it its earlier gain, a higher share of lines carrying a claim a numerical check could test. Before it was taken, an audit found that its significance figure compared two small models within one run, while the registration asked about the drop from the run before, so even a total collapse would have printed "the drop is not significant". An amendment made before the data could steer it fixed the figure, and the same auditor re-audited the fix. Taken once that day, the verdict found the compressed model's share 1.87 points higher, its whole interval above the limit of a 2.0-point fall.
3. Two verdicts whose figures tested something else (August 25, 2026). The earlier verdict on moving that model from Gemma 4 26B to Gemma 4 31B printed the figure of a two-model comparison within one run, though its registered question compared the model's change with a control model's; the verdict holds on that question too. The 4-bit verdict printed the figure the amendment had just fixed, which tests the drop against zero, while the verdict rests on the 2.0-point limit. Both verdicts stand, with notes beside them, and a rule added that day makes every registration name exactly what its significance figure tests.
4. A group of programs folded into another's pass rate (August 29, 2026). A quarter of the places for programs in each run are reserved for kept entries (written by Claude Opus 5 for the small-model lines it chose) that rest on at least three sources, registered as a group of their own. Older pass-rate code folded them into another group, so the report showed 12 passes and 1 failure for a group whose own programs had 8 and 1. No verdict had used the numbers; the code was fixed, and the report stands with its composition disclosed.
5. A recount that could read evidence from after the run (August 29, 2026). The audit reproduced a count of the run's kept entries exactly at the run's end, but a later recount on the live database differed, because entries filed after the run had entered one kept entry's lineage. Since then every re-derivation cuts the data at the run's end.
6. A result a declared ceiling held back (August 29, 2026). This was procedure at work, not an audit catch: one program's single setting moved a computed quantity far enough to look like a first, and a ceiling declared before the program ran parked the result. The session that graded the run (Claude Fable 5, Anthropic) ruled it the source entry's own pre-declared negative outcome, not a first; a review of September 1 to 2 by Claude Fable 5.1 sessions found one row of the program's table to be floating-point rounding and the explanation wrong, and the conclusion stood on the rest.
The later audits, through September 22
| Run | Rows reproduced | Differing | Not re-derivable | Catches |
|---|---|---|---|---|
| August 29 to September 1 | 44 of 49 | 3 | 2 | 3 |
| September 2 to 5 | 122 of 131 | 7 | 2 | 17 |
| Ended September 14 | 138 of 144 | 5 | 1 | 9 |
| September 14 to 19 | 306 of 318 | 6 | ||
| September 14 to 19, second audit | 317 of 334 | 12 | 5 | 7 |
The audit of September 19 was a fresh Claude Fable 5.1 session (Anthropic). A catch and a differing row are different counts: one catch can produce several differing rows, or none.
The first catch of September: a screen that looks for numeric record claims in the small models' lines had reported nothing filed in two runs, but the step that saves a match had never been allowed to write, and the second run's one match was refused (the note on it). The same audit found a benchmark declared frozen eroded by the harness's re-review of old rejections, and a report paragraph quoting another run's figures; both were fixed then.
Six of GPT-6 Astra's seven catches were the first audit's. Its own was code that dropped seven rows from a fixed benchmark, whose verdict stands because both readings clear the same bar (the full account). Its brief had disclosed the topics of the first audit's six fixes, so its independence was of calculation and judgment. A Claude Fable 5.1 session verified every one of its differences against the database on September 22 and fixed what stood. One such audit on one run does not measure how much Claude auditors miss in general.
What this report does not show
The audits set one session's reading of the rows against another session's. The run that ended September 27, 2026 is unaudited. The rows themselves are in the harness's private record, so none of these counts can be recomputed from this site.