Live figures

The loop, live:
each part of Hypnos, what it does, and what its figures count

Hypnos is a research harness, a program whose loop has run since August 11, 2026, with pauses between runs. On graphics cards in the home rack of the owner, the person who runs Hypnos, three small open models read entries that a program picks from a growing notebook of mathematics and write one-line candidate ideas about them; Claude Opus 5, a frontier model made by Anthropic, reads samples of those candidates as the judge and writes the few it keeps into the notebook as entries, where chains of entries build up. In sessions the owner starts by hand, Claude Fable 5.1 (Anthropic) and GPT-6 Astra (OpenAI) read those chains and do the deriving and checking; the layer below them is built to shape what the sessions read, and whether it saves them anything is untested.

How it works explains the mechanism and why it is built this way; the worked case follows one chain from a small model's ten-word line to a theorem in a paper Claude Fable 5.1 wrote.

The figures on this page are from October 11, 2026, 4:25 a.m. EDT; the harness sends a new summary about every ten minutes. The loop is running: 3 of 3 small models are writing candidates.

Figures marked "this run" count from the start of the current run, a period of operation under fixed settings: most are frozen for the run, and a few, such as the filters' thresholds and the judge's model, can change inside it, which the harness tracks. Figures marked "since August 11" count from the first day, August 11, 2026.

What each part does, in the order it runs

Figures marked live come from the latest summary the site received.

  1. 01

    The notebook

    The notebook is the harness's memory: typed entries (mechanisms with what they accept and emit, questions, conjectures, numeric observations, known results) and typed links between them, kept in one database that is the only channel between the harness's parts. Nothing in it is ever deleted.

    What it passes on: one or two entries at a time, which a program picks. The first is drawn by weight from a random walk over the links, recently updated entries weighing more; 65 percent of the time the second is drawn the same way, independently of the first, and 35 percent of the time from the quarter of the notebook farthest from the first in meaning.

    Entries, every type (live)
    3,276
    Links, every kind, including comparison links the walk no longer follows (live)
    14,029
  2. 02

    The small models

    The small models are Gemma 4 31B and Gemma 4 26B, made by Google, and Qwen3 32B. One of them gets the one or two entries the program picked, as short briefs (descriptions cut at 500 characters), and one of five tasks: for a pair, name a non-obvious connection, name a conflict, or propose one small statement a computation could check; for one entry, ask the most basic question nobody has asked, or name the known result it resembles.

    It answers with up to five one-line candidates, each with a stated probability, and is asked to include long shots; most are expected to be poor, since their job is coverage, not correctness. They cannot write to the notebook.

    Small models writing now (live)
    3 of 3
    Candidates written this run (live)
    37,772
    Candidates written since August 11 (live)
    645,590
  3. 03

    The filters and queues

    Code, with no model involved, runs every candidate through a fixed order: a numeric check of statements about zeros of the zeta function; removal of near-duplicates; a score (novelty against the notebook, structural fit, testability, source trust); a floor on that score; then a queue of at most 400 per small model. When a queue overflows, its lowest-scoring candidates are set aside, never deleted.

    The filters remove almost nothing: in the run of September 19 to 27, 2026, 99.6 percent of candidates passed them, and the queue cap set aside most of those. What goes on to the judge: the queued candidates.

    Set aside by a full queue, since August 11 (live)
    611,974 (95 percent of all candidates written)
  4. 04

    The judge

    Claude Opus 5 (Anthropic) reads candidates in batches of 27 under fixed rules that expect almost everything to be "slop" and keep only the rare item that is non-obvious, structurally grounded and usable by a later session. It is shown each candidate's stated probability, which does not decide what it reads.

    From August 11 to September 27, 2026, it kept about 3 in 100 of the candidates it read; in the run that began that day, under stricter rules, about 1 in 100. Those rules let it keep an item only when four required marks of five are yes, and an exception must name the mark it overrides; a planted-error test showed they also reject more error-free items.

    Candidates read this run (live)
    2,104
    Kept this run (live)
    2 (0.1 percent of those read)
  5. 05

    Kept entries

    A kept entry is the judge's own writing, not the small model's line: a name, a type and a description a later session can act on, a median 11 times the line's length, sometimes with a computation to run. It enters the notebook at the lowest trust level, linked to the entries the line was about.

    A new entry has full recency weight in the walk, so later picks near its sources reach it, and over weeks entries linked to their sources form the chains of entries the sessions read. An entry that carries a computation to run becomes eligible for the execution step.

    Kept since August 11 (live)
    863
  6. 06

    The execution step

    A kept entry that carries a computation to run can become a program. Separate calls to the judge's model, Claude Opus 5, write the checks (one must reproduce a known value, since the same model writes both), then the program. It runs without network access, using verified routines built on the four instruments and standard numerical libraries. Code grades it: pass, fail, or could not measure, a verdict on the program rather than the idea.

    The result is filed as an entry linked to the idea it tested; trust does not move. By September 29, 2026 there had been 339 such programs, 124 of them passed; two of the four mathematics papers cite measurements made this way.

  7. 07

    Consolidation and the mirror

    After every 12 new entries, code asks the judge's model, Claude Opus 5, to compress recent growth, and writes back themes linked to their members, proposed merges, cross-links and questions about missing bridges. After every 40, a second pass, the mirror, retells the notebook as four essays under a rule to assert nothing new, so that a pattern hidden in one telling may show in another.

    Code parks entries under consolidation's rules: when the notebook holds more than a set number of speculative entries, it sets aside the least recently touched ones that no kept entry cites. A parked entry stays in the notebook, at half weight in the walk.

  8. 08

    Checks on the loop

    Five of each batch of 27 come from what was set aside, from removed duplicates and from the judge's own rejections at least eight batches old; the prompt does not label them, though the judge sees every item's score. These blind draws measure what each of those steps loses.

    From August 11 to September 27, 2026, the blind draws were kept at least as often as candidates from the queues, so no step of the layer has yet been shown to select better than a blind draw. The four rows below add up to 23,382 items read, against a total of 22,464 read by the judge over the same period; the sources do not reconcile the two.

    From the queues
    19,085 read, 591 kept, 3.1 percent
    Blind draws from what was set aside
    2,591 read, 87 kept, 3.4 percent
    Blind draws from removed duplicates
    854 read, 35 kept, 4.1 percent
    The judge's own rejections, redrawn
    852 read, 49 kept, 5.8 percent
  9. 09

    Trust levels

    Every kept entry enters at the lowest of four trust levels (speculative, numerically supported, proved informally, proved formally in Lean), and nothing in the running harness raises it; the code that could, on a recorded check, is called by nothing but a test. Since trust is constant in practice, what the walk's weights track is recency: an entry's weight halves every 72 hours after it was last updated.

    A separate pipeline for formal proofs in Lean was built outside the running loop; by August 13, 2026 it had checked three formal statements, and the loop sends it nothing. One entry's life in the notebook is followed on How an entry changes.

  10. 10

    The inlets

    The owner's seeds are his own speculative notes, kept permanently at the lowest trust; the small models are told they are "raw aim, never evidence". He meant them as "sprinkles on a cupcake"; when a measurement on August 13, 2026 showed lenses made from them (a lens is a stored way of looking at the problem that tilts where the walk starts) taking 86.5 percent of search attention, the fix was a guaranteed floor for every kind of lens, and lenses the notebook mints itself.

    Known results are filed as entries to check against; they are not themselves lenses. Ideas from the public web are filed as entries with their authors' names attached.

    The owner's seeds (live)
    25
  11. 11

    The target problem

    The notebook is on mathematics near the Riemann hypothesis. It was seeded on August 11, 2026 with four ways of looking at the hypothesis (an operator whose spectrum would be the zeros, Weil's positivity criterion, random matrices, function fields), four classical bridges between fields, the owner's starting ideas and four computational instruments, each checked against published values before any research could trust it.

    The problem was chosen as a load test for the harness, not as a success criterion, and the owner does not expect the small models to solve it.

  12. 12

    The sessions

    The loop hands the paper-writing models nothing. The owner starts a session with Claude Fable 5.1 or GPT-6 Astra and asks it to find something worth proving in the harness's work; the other company's model reviews the paper.

    Four mathematics manuscripts came from such sessions; as of October 3, 2026, no human mathematician has read any of them. Only one question grew through the notebook from a small model's line; one paper took five notebook entries as motivation; one note answers two questions J. J. Gil posed in a 2026 paper; one came from sessions on the OpenAI forced Navier-Stokes blow-up manuscript, whose material is not in the notebook. How it works gives each origin in full.

Where to go next

What Hypnos is doing now, in figures
The funnel from candidates written to programs run, live.
Entries kept recently
The latest entries the judge kept, with titles written by a small model.
The home page
What came of the work, and how to take part.