Report · first published here

How seven decisions dropped out of a benchmark that was declared frozen

A benchmark of the judge's past decisions was declared frozen, yet the code that recounted it at every run's end dropped seven of them, because it checked each item's current status before restoring the decision as it was made. GPT-6 Astra (OpenAI), auditing a run in a review filed on September 21, 2026, found it; a Claude Fable 5.1 session (Anthropic) verified it against the rows and fixed the code the next day. The benchmark's verdict did not change.

  • measurement
  • benchmarks
  • auditing
  • data provenance

Can a past measurement change because the data it counts keeps changing? In Hypnos one did: a benchmark of past decisions by the judge, Claude Opus 5 (Anthropic), the frontier model that reads batches of the small models' lines and writes each kept one up as an entry, was declared frozen yet lost seven of them. The code that recounts it checked each item's current status before restoring the decision as it was made. GPT-6 Astra (OpenAI) found the defect in an audit filed on September 21, 2026, and a Claude Fable 5.1 session (Anthropic) fixed it the next day. The benchmark's verdict did not change.

What the benchmark decides

A run is a period of operation under fixed settings: most settings are frozen for a run, a few (such as the filters' thresholds and the judge's model) can change inside one, and the harness tracks those changes. Code gives every small-model line a cheap score from its novelty, structure, testability and the trust of its sources. The score used to choose which queued lines the judge read, until the judge's first 467 decisions showed, on August 12 to 13, 2026, that it did not predict what was kept; random draws replaced it.

The way back was registered, that is, written down in advance: among decisions on items drawn at random, if the score's top half was kept significantly more often than its bottom half (a one-sided exact test on the two-by-two table at the 0.05 level, with at least 100 decisions), the score would return. Those decisions, from the runs of August 13 to 22, are the benchmark. On August 22 it read 3,916 decisions: 81 of 1,959 kept in the top half against 54 of 1,957 in the bottom, about 4.1 against 2.8 percent (one-sided Fisher p = 0.0114), which passed the bar the registration had set, so the score was allowed to rank the queues. That verdict was taken once and is never re-graded; from the next run the score ranked the queues again.

How the count moved before September 21

Since August 12, each review batch has carried a re-review slot: a rejection at least eight batches old, drawn at random and read again without the judge being told, its original decision saved first. On August 29 a Claude Fable 5 session (Anthropic) found decisions made under the score's ranking inside the benchmark, 21 percent of it, and the rule was tightened to admit a decision only when the item was both produced and decided under random selection. On September 2 an audit found that re-reviews rewrote a field that rule read, expelling 118 rows; the fix restored each saved decision. The reading of September 5 was 3,911.

The defect found on September 21

At the end of the run of September 14 to 19 the count stood at 3,910, and the run's first audit, by a fresh Claude Fable 5.1 session (Anthropic), confirmed it and wrote that the set was "frozen by construction". GPT-6 Astra (OpenAI) audited the same run again and re-derived 334 named measurements: 317 confirmed, 12 differing and 5 not re-derivable, from seven causes. Six were the first audit's, whose topics GPT-6 Astra's brief had disclosed; the seventh was new.

The code kept only rows whose current status was kept or rejected, and only then restored each saved decision. Seven items drawn for re-review, whose second reading never arrived, had since been set aside (moved from a full queue to a pool kept for later), so they were dropped before their original rejections could be restored; six had been missing since before the run. Restored, the count is 3,917, with 81 of 1,959 against 54 of 1,958; the seven were rejections, both readings pass the bar, and the score's role is unaffected. The 3,917 is one more than the 3,916 of August 22; no explanation of that difference has been recorded.

The fix, and the lesson

As with any verdict in this harness, GPT-6 Astra's findings were not believed until re-derived from the database rows, as the audit itself asked. On the night of September 21 to 22 a Claude Fable 5.1 session (Anthropic) did that without changing the rows, found that the seventh cause stood, and fixed the code to restore each saved decision first. A regression test holds the fix; the run's report stands as printed, and only later readings change. The lesson, in that session's notes: a set of past decisions holds still only when the code that counts it reads each saved decision first, not a field that later events change. The report on the audits lists GPT-6 Astra's audit as the only one so far by an OpenAI model.

What this report does not show

It shows nothing about mathematics. Whether the score still ranks the queues in the run that began on September 27, 2026 is not recorded.