Live figures
What Hypnos is doing now, in figures
Hypnos is a research harness whose loop has run since August 11, 2026, with pauses between runs. In the loop, three small open AI models write one-line candidate ideas about entries that a program hands them from a growing notebook of mathematics; they run at home, on the graphics cards of the owner, the person who runs Hypnos. Claude Opus 5, a frontier model made by Anthropic, is the judge: it reads samples of those lines and writes the few it keeps into the notebook as new entries. In sessions the owner starts by hand, two other frontier models, Claude Fable 5.1 (Anthropic) and GPT-6 Astra (OpenAI), derive and check new work; the loop is built to shape what the sessions read, and whether it saves them anything is untested. How it works explains each step.
This page shows what the loop has done since August 11, 2026 and what it was doing at the last update. The harness normally sends the site a summary about every ten minutes; its own log stays private.
The latest figures reached this site on .
Live totals since August 11, 2026
645,590 candidate lines the small models wrote, most of them never read: each model's queue holds 400, and the rest are set aside, never deleted.
30,987 read by the judge, in batches of 27, under written rules that expect almost everything to be poor.
863 kept, each written up by the judge as a notebook entry of its own, linked to the entries its line was about.
The bars are on a logarithmic scale, so the small counts stay visible.
Computed from these live totals: the judge has read about one in 21 of the candidates written and kept 2.8 in 100 of those it read, so about one in 748 of all candidates written has become a notebook entry.
1,263.2 hours on the loop's clock, which stops between runs and while every small model's server is unreachable, so it shows less than the time elapsed. Over those hours the small models wrote about 511 candidate lines an hour.
The loop's state
Running, paused or between runs
The loop is running: the small models are writing candidate lines, and the judge reads a batch about once an hour of the loop's clock.
Run by run
The same figures, one run at a time
A run is a period of operation under fixed settings; most settings are frozen for a run, a few (such as the filters' thresholds and the judge's model) can change inside one, and the harness tracks those changes. Changes to the harness are made between runs, and the harness measures itself one run at a time.
| Last figures received | Hours on the clock | Read by the judge | Kept | Kept per 100 read |
|---|---|---|---|---|
the run in progress | 78.7 | 2,104 | 2 | 0.1 |
| 255.5 | 6,690 | 99 | 1.5 | |
| 177.9 | 4,697 | 149 | 3.2 | |
| 121.7 | 3,172 | 106 | 3.3 | |
figures stopped early | 53.9 | 1,346 | 52 | 3.9 |
| 83.6 | 2,149 | 72 | 3.4 | |
| 79.6 | 2,122 | 70 | 3.3 |
Each run is dated by the last figures the site received for it, not by the day it began or ended; a run whose figures stopped before it ended is marked. The table starts with the earliest run whose figures the site kept. Hours are on the loop's clock, which stops while the loop is paused.
What changed between these runs
The judge's rules became stricter on September 27, 2026: in the rows dated after that day, the judge kept about 1 in 100 of what it read, against 3 to 4 in 100 before (each row gives its counts). It now marks each candidate 0 or 1 on five counts and may keep it only when four are 1 (grounded in its source, states a mechanism, not trivial, honest about its scope), and an exception must name the mark it overrides. In a test that day on 24 candidates with planted errors, the old rules kept 11 and the new rules 1; the new rules also rejected three more of the same candidates written without the errors.
The random walk that weights which entries the small models are handed changed on September 19, 2026: it now follows only links that record an event, such as a kept entry's link to its sources, and no longer the links a program writes for every technique it compares with a problem, pass or fail (37 percent of all links on September 27). A comparison set up in advance could not tell the two ways apart, so the change rests on an argument: such links say nothing about the entries they join.
The notebook
How big the notebook is
The notebook is the harness's memory, and nothing in it is ever deleted. Any two of its entries could be handed to a small model together, and the number of such pairs grows with the square of the number of entries: far more than a frontier model could read, which is why small models walk them all day. How a program picks what the small models read
Live figures
3,276 entries
14,029 links between them
5,364,450 possible pairs of entries
Elsewhere
Where to look next
- The loop, live: each part of the loop in the order it runs, with its live figure.
- Entries kept recently: the latest entries the judge kept, with titles and summaries Gemma 4 26B, one of the small models, writes for this site.
- One chain from end to end: from a small model's ten-word line on August 12, 2026 to a theorem Claude Fable 5.1 proved on September 29, with who wrote each step.
- How to check the papers and take part, on the home page.