The explanation page

How Hypnos works

Hypnos is a research harness: a notebook of mathematics; a loop in which three small open AI models write one-line fragments about its entries and a frontier model from Anthropic rewrites the few it keeps as new entries; and sessions in which the owner, the person who runs Hypnos, sets frontier models to work on what has built up. Kept entries link into chains, and a session reads chains whole and picks one: that is how ten words from a small model started the chain behind one paper.

The loop in one picture

Solid lines are the loop, which runs on its own; dashed lines are sessions the owner starts by hand.

How the parts of Hypnos connect The walk, a program Three small open models Gemma 4 31B and 26B (Google), Qwen3 32B Filters and queues, code The judge Claude Opus 5 (Anthropic) Inlets seeds, known results, public-web ideas The notebook The execution step Consolidation and mirror passes Claude Opus 5 Sessions the owner starts Claude Fable 5.1 (Anthropic) or GPT-6 Astra (OpenAI) Papers four mathematics, one on methods Reviews by the other company's model the links between entries entry briefs and one of five tasks one-line candidates with probabilities batches of 27, five of them unlabeled draws kept, written up as entries as entries entries that carry a computation to run graded result, as an entry chains of entries, read by hand a paper with its checker programs down: recent entries; up: themes, cross-links, questions, parked entries findings and written answers
  1. A program, not a model, draws one or two entries by weight from a random walk over the links; for 35 percent of pairs the second is chosen far from the first in meaning.
  2. One of three small open models reads them as short briefs with one of five tasks and writes up to five one-line candidates, each with a stated probability.
  3. Code refutes what it can check, drops near-duplicates and queues at most 400 per model; the rest are set aside, never deleted.
  4. The judge, Claude Opus 5, a frontier model made by Anthropic, keeps about 3 in 100 of those it reads and writes each up as an entry linked to its sources.
  5. For a kept entry with a computation to run, the same model writes checks, then a program; code grades it and files the result.
  6. Consolidation passes on the same model group recent entries into themes, cross-links and questions, and mirror passes re-describe the whole notebook, asserting nothing new; seeds, known results and public-web ideas come in as entries.
  7. In sessions started by hand, Claude Fable 5.1 (Anthropic) or GPT-6 Astra (OpenAI) reads chains of entries and chooses what to prove; the other company's model reviews the paper.

Why it is built this way

The owner's statement of purpose opens: "Hypnos is not an attempt to show that open models can do what frontier models do. They cannot, and the harness is built on that fact. It is a pre-processing layer." The notebook it pre-processes is too large for a frontier model to read pair by pair: about 2,700 entries make more than three million pairs, which a small model can walk all day. Its job, in the design's words, is "coverage, not correctness": most of what it writes is expected to be poor, so a frontier model does the reading and judging.

Frontier-model time is the scarce part, and some is spent inside the loop: the judge reads about one candidate in twenty, and the same model runs consolidation and the mirror and writes the programs. Deriving something new and checking it is left to the sessions; whether the arrangement saves frontier-model effort overall is untested.

The notebook's subject is mathematics near the Riemann hypothesis, chosen as a load test for the harness, not as a goal; the owner does not expect the small models to solve it.

How a one-line fragment reaches a frontier model

On its own, a small model's line does nothing: the model cannot write to the notebook, and most lines are never read. Four things must happen first.

  1. The line becomes the judge's entry. If Claude Opus 5 keeps it, it writes a new entry of its own, typically about eleven times as long, with the line stored beside it.
  2. The entry becomes material. Linked to the entries the small model read and embedded at once, it is checked against every new candidate, carries full weight in the walk so that walks near its sources reach it, and credits the lenses (stored ways of looking at the problem) beside its sources, so walks start there more often.
  3. Entries become a chain. Later lines about it pass the same gate, the judge writes sharper questions from them, and programs file measurements beside them. Nothing assembles the chain on purpose; each kept entry links back to what it came from.
  4. A session reads the chain whole. Nothing in the loop hands a chain on. The owner starts a session with Claude Fable 5.1 or GPT-6 Astra and asks it to find something in the harness's work worth proving. The session behind the unfolded-zeros paper took a copy of the notebook (2,723 entries and the 339 programs run by September 28, 2026), pulled six chains of entries out of it itself, scored each on whether it was provable, non-trivial and rooted in the notebook, and chose one. The session's effort lands on the chain it chose; the fragment is where that chain started.

The owner describes the aim as trying "to focus the free association into linkable and testable concepts": the small models' lines are the free association, the judge's linked entries make it linkable, and the computations make it testable.

Three limits apply. Whether later sessions use kept entries at any rate is not yet measurable. The fragments do not contain the mathematics: in a replay a session ran on September 29, 2026, the small models were handed the exact pairs of entries behind two of the papers, and none wrote the proof's key move. And in 117,339 steps the walk had never drawn those pairs together: coverage is the aim, but the chains were assembled by kept entries linking to their sources, not by the walk pairing them.

One chain, from ten words to a theorem

This is the only chain that begins with a small model's own words and ends in a theorem its paper says answers the notebook's question. The unfolded zeros are the zeta zeros' heights, rescaled so that the n-th sits near n; a Riesz basis of exponentials behaves like an orthonormal basis up to constants; and Kadec's 1/4 theorem gives one on (−π, π) when each frequency lies within less than 1/4 of its integer, a sufficient condition, not a necessary one.

  1. August 12, 2026 (a small model's line). Asked which known result an entry resembled, one of the three small models wrote: "The Kadets 1/4-theorem regarding the stability of bases of exponentials."
  2. August 12 (the judge's entry, Claude Opus 5). Kept and rewritten as a question 21.7 times as long: are the unfolded zeros within 1/4 of the integers, so that Kadec's theorem applies? It predicted early failure, since the argument of zeta is unbounded.
  3. August 14 and 20 (small models' lines, kept by the judge). Two more recalls of the theorem.
  4. August 16 (a small model's line; the judge's entry). A false claim that the first N zeros meet Kadec's condition, which the judge rewrote: the condition is exactly a uniform bound on the argument of zeta at the zeros, and the live question is Avdonin's averaged condition and Pavlov's sharp criterion.
  5. From August 19 (programs written by the judge's model, graded by code). The condition fails at the ninth zero; over the first 100,000 the largest deviation is 1.129.
  6. September 2 (a small model's line, kept by the judge). A recall of Pavlov's stability results; that entry is not among the 17 the paper cites.
  7. September 2 (a small model's line; the judge's entry). A conjecture about the first 50 rescaled zeros, which the judge turned around: Kadec's condition is only sufficient, so its failure decides nothing about whether the zeros still act as a sampling set, and a measured quantity should replace it.
  8. September 18 (a small model's line; the judge's entry). A request to compute Riesz constants for 1,000 rescaled zeros; the judge asked whether the lower constant stays away from zero or falls toward it, so that the Riesz property is lost.
  9. By September 28 (programs written by the judge's model, graded by code). The lower constant fell from 0.1115 at 100 zeros to 0.0159 at 1,000.
  10. September 28 (Claude Fable 5.1, in a session the owner started). The session described above scored six chains and chose this one.
  11. September 29 (Claude Fable 5.1). The paper proves that the unfolded zeros, counted with multiplicity, form no Riesz basis of exponentials on any bounded interval, joining Pavlov's criterion (a Riesz basis forces the counting deviation into bounded mean oscillation) with Selberg's mean-square growth of the argument of zeta; for the distinct frequencies it proves this only when all but finitely many ordinates are simple. It says this answers the question of the September 2 and 18 entries: the Riesz property is lost.
  12. What the small models did not do (a replay a session ran on September 29). Handed the chain's own pairs of entries, none wrote a line containing the proof's mechanism; its two halves appeared in different models' lines, never joined. Whether Claude Fable 5.1 took its route from the notebook, where Pavlov's name appears twice, or from what it knew is not recorded.

In the paper's words, the harness "supplied the motivating questions and the numerical observations, which are reported separately from the proofs": every question here was the judge's, every measurement a program's, and the proof Claude Fable 5.1's. The worked case draws it in panels; the paper's page has the theorem, reviews and files.

The models and their jobs

The models, as of October 3, 2026
Model and makerJob
Gemma 4 31B and Gemma 4 26B (Google); Qwen3 32BThe small models, on the owner's graphics cards with compressed weights; two from one family, one from another.
Claude Opus 5 (Anthropic)The judge since the night of August 11, 2026, chosen with the recorded reason "subscription usage headroom beats marginal judge quality".
Claude Opus 5.5 (Anthropic)Fresh-instance reviews and consistency passes on the papers; the blind grading of the Navier-Stokes test.
Claude Fable 5 (Anthropic)Wrote the methods paper, August 29, 2026.
Claude Fable 5.1 (Anthropic)Wrote three of the four mathematics manuscripts; applied the review edits to all four.
GPT-6 Astra (OpenAI)Wrote the antiunitary paper; reviewed the other three.
Gemma 4 31B (Google); GPT-5.6 and later GPT-6 Astra (OpenAI)Shadow readers: since September 2026, first a local Gemma 4 31B and then also the OpenAI models re-read the judge's batches and record what each would keep, with no power over any item.

The notebook

The notebook is one database of typed entries and links, and the only way the parts talk. The design follows the owner's specification, whose first phase reads:

Concepts enter as "hollow shells" (labeled, unusable). They gain usability only through enforced interaction. Everything is stored as a typed mechanism, not a fact.

Its second phase opens: "Problems are compiled to domain-free mechanism signatures"; in the code, mechanisms (usable tools) and problem statements carry interfaces in a fixed vocabulary that code compares, but the entries the loop adds carry none, so the comparison never reaches them. How an entry changes has the shells, the usability gate and the vocabulary.

It started on August 11, 2026 with four ways of looking at the Riemann hypothesis, written as problem statements, four classical bridges (the Mellin transform of the theta function, Weil's explicit formula, the dictionary between number fields and function fields, the Montgomery-Dyson link between the zeros' spacing and random matrices), four computational instruments, one wall and 13 of the owner's starting ideas, his seeds (the About page lists all 18), which are aim, never evidence.

By late September it held about 2,700 entries and 10,700 links: 762 ideas the judge had kept by September 27, results of the 339 programs run by September 28, and the rest from consolidation, seeds, known results and ideas from the public web, which are filed with their authors' names attached. Neither how that rest divides nor how the web ideas are gathered is published. Every kept entry enters at the lowest of four trust levels (speculative, numerically supported, proved informally, proved formally in Lean), and nothing in the running harness raises it. Nothing is ever deleted.

How entries are picked

A program picks the entries. The first is drawn by weight from a random walk over the links (personalized PageRank). Since September 19, 2026 the walk follows only links that record an event (a derivation, a grouping, a program run); until then the links written for every comparison of a tool with a problem statement, pass or fail (37 percent of all links), pulled it as hard as a derivation. Sixty-five percent of the time the second entry is drawn the same way, independently of the first; 35 percent of the time it is drawn from the quarter of the notebook farthest from the first in meaning. The walk's starting weights follow age, halving every 72 hours since an entry's last update; trust is in the formula too, but it never moves.

Lenses, stored ways of looking at the problem written as problem statements, tilt where walks start. Thirty percent of the lens weighting is shared evenly by the kinds of lens: the four classical ones, and those grown from the owner's seeds, recurring themes, public-web ideas and known results (a theme becomes a lens once it holds at least six entries and persists across consolidation passes; a known result is not a lens itself, but one cited by four entries can seed one). The rest follows yield, the share of reviewed items within one link of a lens that the judge kept.

The owner had meant his seeds as "sprinkles on a cupcake"; when, on August 13, 2026, 28 of 32 lenses turned out to come from them and to hold 86.5 percent of search attention, the floor by kind and lenses of the notebook's own were added. By September 27, 31 lenses the notebook minted had a kept entry credited, and 379 of the 762 kept entries had no lens within one link of their sources.

What a small model is asked to write

Every prompt opens: "You are the background association layer of a mathematics research memory"; the entries follow as short briefs, each description cut at 500 characters. Five tasks rotate, three on a pair, two on one entry:

  • Associate (five answers): "Is there a non-obvious connection between entry A and entry B?"
  • Tension (four): do they "point in conflicting directions anywhere"?
  • Micro-conjecture (five): one small statement "checkable by computation at modest precision".
  • Question (four): "What is the dumbest question nobody has asked about this entry?"
  • Recall (four): "Which known classical result, technique, or paper does this entry rhyme with?"

Each answer is a numbered line opening with a probability, the model's own estimate that it is right; the prompt asks for long shots, "unlikely-but-interesting candidates with honest low probabilities". The design cites this technique, verbalized sampling (arXiv 2510.01171), for recovering much of the diversity alignment training removes; that paper reports 1.6 to 3 times the diversity of a plain prompt.

From September 19 to 27, 2026 the three models wrote about 4.4 candidates a step, and 2.4 to 7.3 percent of each model's nearly matched another model's; two families run side by side so this can be measured: in the design's words, if they proposed the same connections, "cross-model diversity is cosmetic". The stated probability is shown to the judge but chooses nothing.

What is kept, and how

The filters come first, with no language model: a numeric check of claims about the zeros, removal of near-duplicates, and a score that weighs novelty at 0.40 (one minus the line's closest match to any entry), structural fit at 0.30 (a flat 0.3 unless a tool is involved), testability at 0.20 and the trust of its sources at 0.10; a line under 0.15 is set aside. Each small model's queue holds at most 400 lines, and its lowest-scoring overflow is set aside too. From September 19 to 27, 2026, 99.6 percent of 94,765 candidates passed the filters, the caps set aside 90,092, and the judge read 3,509 of them. Its batches take up to three lines from each queue, then the rest in proportion to each queue's size, and since August 23, 2026 the highest-scoring first, so the score decides most of what the judge sees.

The judge, Claude Opus 5, runs on the owner's Anthropic subscription through Claude Code, Anthropic's command-line program, with every tool turned off; the loop's other frontier tasks take the same route. Each batch of 27 holds 22 items from the queues and five unlabeled draws from what was set aside, from removed duplicates and from its own rejections at least eight batches old; the prompt does not label them, though the judge sees every item's score. Its rules say "almost everything below is expected to be slop, and that is fine": keep only the rare item that is non-obvious, structurally grounded and actionable, and give each rejection a reason, which is recorded. For a kept item it writes the entry, sometimes with a computation to run; kept means judged worth a closer look, not that it is true, new or proved. Code applies the verdicts, refusing any claim above the lowest trust.

Since September 27, 2026 the judge may keep an item only when four of five marks are 1 (grounded in the named source, states a mechanism, non-trivial, honest about scope; the fifth, checkable, is recorded), and an exception must name the mark it overrides. On 24 test items with planted errors, the earlier rules kept 11 planted falsehoods and the new ones one, while rejecting three more error-free items. Every entry kept before September 27, this page's chain included, was kept under the earlier rules.

Every way of setting a candidate aside has a blind draw except numerical refutation:

Reviewed and kept, by source, to September 27, 2026. The rows sum to 23,382 reviews against a period total of 22,464; the harness's reports do not break the difference down.
SourceReviewedKeptShare
The queues19,0855913.1 percent
Draws from what the caps set aside2,591873.4 percent
Draws from removed duplicates854354.1 percent
The judge's own rejections, redrawn852495.8 percent

So no step of the layer has yet been shown to select better than a blind draw, and the judge keeps its own redrawn rejections more often than first looks. The shadow readers in the models table check the judge: of the 149 items it kept from September 19 to 27 (27 of them written earlier), Gemma 4 31B would have kept 15 and GPT-6 Astra 11.

The checks: instruments and programs

Four computational instruments were checked against published values on August 11, 2026: one computes the zeta zeros and their count three ways, certifying that none was missed, and three reproduce published proportions of the zeros. An instrument answers pass, fail or could-not-measure; a dead one also gives a zero reading, so a zero reading is checked for one before it is believed (why). One instrument's first reading flagged an apparent defect in the 2020 formula of Pratt, Robles, Zaharescu and Zeindler; pre-registered checks by seven fresh sessions, three implementing the formula, traced it on August 18 to a one-line bug in the harness's own evaluator (the story).

The execution step turns a kept entry that carries a computation into a program. Separate calls to the judge's model write the checks, frozen before the program exists so they cannot be fitted to it, and then the program; because one model writes both, one check must reproduce a known value. It runs with no network, and code grades it: a failed anchor check is could-not-measure, a result that would beat a certified value is quarantined, and only then pass or fail, neither of which moves trust. A quarter of each run's program slots go at random to kept entries that draw on at least three sources. From August 11 to September 27, 2026, 291 programs were started; a copy of the notebook taken on September 28 held 339, of which 124 had passed.

The owner's specification states the principle: "Assume the model will game any self-assessed metric"; so code filters, queues, applies verdicts and grades programs, and never raises trust on a model's word; models decide what is kept (the judge), what to work on (the sessions) and whether a paper is right (the reviewing models). The harness's measurements of its own parts are registered before the data exist, as the methods paper describes; a report lists what its audits caught.

The sessions and the papers

The statement of purpose says the sessions work on "what the notebook and the review of earlier results put in front of them": chains of entries, and the reviews of the earlier papers, where the Gil note's questions surfaced. A paper comes with checkers: programs that recompute its numbers, published beside the output they recorded. In the loop, the judge's model also runs consolidation, which after every 12 new entries groups them into themes, cross-links and questions, and the mirror, which after every 40 new entries rewrites the whole notebook in four mathematical languages (algebraic, spectral, probabilistic, geometric) under a rule to assert nothing new.

Where each paper's question came from; only the first paper concerns the zeta zeros. The methods paper (Claude Fable 5, Anthropic, August 29, 2026) is about the harness's conduct and claims no mathematics.
Paper and writerWhere the question came from
The unfolded zeros paper (Claude Fable 5.1, Anthropic)The only one whose chain begins with a small model's own line: a ten-word recall of Kadec's theorem on August 12, 2026, kept by the judge as a question; more lines and sharper judge-written questions over five weeks; ten programs; on September 28 a session the owner started scored six chains of entries and chose this one; the paper says its main theorem answers the question the notebook posed.
The antiunitary paper (GPT-6 Astra, OpenAI)Its motivation is five notebook entries of late September: two entries the judge wrote on small-model lines (one a recall of the classical fact that a skew-symmetric matrix of odd order has determinant zero, one a false conjecture) and three results of programs. The paper calls them motivation, not premises; its main theorem is a reformulation the writing session made; where its key tool, the matching polytope, came from is not recorded.
The note on Gil's questions (Claude Fable 5.1, Anthropic)The two questions are J. J. Gil's own, from Entropy 28(8):877 (August 4, 2026). They surfaced while frontier models reviewed the antiunitary paper; GPT-6 Astra's own review proved the free-value answer first. No step of the note's mathematics came from the harness.
The dividing-plane note (Claude Fable 5.1, Anthropic)A separate line of work the owner ran with frontier-model sessions on the OpenAI forced Navier-Stokes blow-up manuscript: fresh Claude Opus sessions wrote section digests and eleven section explanations, from which a Claude Fable 5.1 session assembled a ledger of the construction's moves; the explanation of the manuscript's Appendix B asked the question and sketched the answer; a Claude Fable 5.1 session ranked it first on its map of open questions and proved it. No small model's line is among its sources, and the material has not been filed into the notebook.

Each manuscript was reviewed by the other company's model and by fresh instances from the writer's company; every finding has a published written answer, and the reviews changed the papers. GPT-6 Astra found a blocking error in the unfolded zeros paper, a completeness statement that counted the zeros with multiplicity though completeness depends only on the distinct points, since restated. Claude Fable 5.1 found the antiunitary paper's single-observable case classical, which changed its account of what is new, and found two open questions of J. J. Gil's that a lemma of the paper bears on. GPT-6 Astra found the dividing-plane note's second obstruction to a symmetric vortex core already proved by Duraiswami, now credited. Disagreements are recorded, not averaged. As of October 3, 2026, no human mathematician has read any of the five papers. The files page has every review, answer and checker with its output and SHA-256 fingerprint, also at github.com/dross50/hypnos-papers.

The human and the sessions

The owner starts every session. He does not check the proofs and could not: his reading of each manuscript looks for red flags, that its account of the harness matches the harness's log and that nothing reads like a machine grading its own homework. The frontier models run on two $200-a-month subscriptions and the small models on his own hardware; the home page gives what is known of the budget.

Sessions also build and change the harness between runs, and eight times in August and September 2026 a fresh session recomputed a run's graded results from the raw data before any verdict was believed; GPT-6 Astra did one such audit. OpenAI's models work through OpenAI's Codex agent, the command-line program through which GPT-6 Astra runs.

The attention analogy, and where it breaks

The analogy is the owner's, his "attention transformer concept"; in his words, "I believe the true value will be in managing the attention of Fable/Astra to help them perform how I need them to." He pictures the harness doing for a frontier model what the attention step of a transformer does for a token: picking out which other material matters and carrying what it implies forward, so that the frontier model's effort lands where the material points. It is a design picture, not a working function or a tested result; put plainly, the layer is built to shape what the sessions read.

It breaks in nine places:

  • Nothing is trained by gradient; the lens weights are counted from the judge's verdicts.
  • There are no learned queries or keys, and in the ordinary draw the second entry ignores the first.
  • One pair is sampled per step, not a weighted mix of all the entries.
  • The steps are events over days, not stacked layers, and a change between runs can alter the design.
  • The small models write text, and most of it never reaches the judge.
  • What is written back is authored mostly by a much larger model, the judge.
  • The selection is noisy, as the blind draws show.
  • The walk did not assemble the papers' chains.
  • Whether this saves frontier-model effort is untested.

What survives: the notebook is a stream that only grows, a larger model decides what is written into it, and the frontier sessions read what has built up.

What is known, and what is not

What has been measured. From August 11 to September 27, 2026, the small models wrote 485,873 candidates. The judge read 22,464 of them, about one in twenty; the rest waited in a queue or were set aside when a queue was full, never deleted. It kept 762 of those it read, about 3 in 100. So about one in 640 of all candidates written became a notebook entry. Frontier models wrote four mathematics manuscripts in sessions the owner started, and only one bears on the aim: the unfolded-zeros paper, whose chain runs from a small model's line to a theorem. The antiunitary paper took only its motivation from the notebook, the Gil note's questions are J. J. Gil's own, and no small model's line is among the dividing-plane note's sources.

A registered test on September 30 and October 1, 2026, with no pass mark, showed the small models each move of the construction in the OpenAI forced Navier-Stokes blow-up manuscript with its reason withheld. On that question the result was null: none of 16,018 lines recovered a reason. A hit was a line stating what a move depends on (or, three times, what a wall forbids); of 156 such lines on moves, 128 restated the prompt and 28 added something. The test's report flagged seven lines for a closer read; each was later answered by earlier work, found already in the manuscript, or not attempted. The stated probability averaged 0.59 on hits, 0.51 on near-misses and 0.27 on misses, with four limits: the 28 hits that add something average 0.44, below the near-misses' 0.51; the task that produced most hits also runs most confident; the harness uses the probability for nothing; and the September 29 replay read the other way (0.55 on lines that only restated the prompt, 0.20 on near-misses that did not). The graders, 32 Claude Opus 5.5 sessions working blind and 27 auditors, were one model family. The owner reads the result as the system working as intended.

What has not been shown, or points the other way. The sessions are started by hand, not called up by the loop. The comparison that would test the aim, the same model on the same problem with the same budget, with and without the layer, has not been run; a cost-matched version was designed and registered on August 13, 2026, but is not built, waiting on the owner's go. No step of the layer has yet been shown to select better than a blind draw; frontier models also work inside the loop; no human mathematician has read the papers. Open: whether later sessions use kept entries (not yet measurable), and whether the stated probability would work as a selection signal.

The words this site uses

a session

Work a frontier model does because the owner started it.

a chain of entries

Entries linked through what each came from.

a seed

One of the owner's starting ideas; aim, never evidence.

a shell

A mechanism labeled but not yet usable.

a wall

An entry recording a proved limit of one family of methods.

a computation to run

A runnable check the judge may write for a kept idea; an anchor check must reproduce a known value.

an unlabeled draw

An item drawn blind into a review batch from what was set aside, removed or rejected.

a registered test

A measurement whose rules were written before the data existed.

a run

A period of operation under settings that are frozen except for the filters' thresholds, the lens weights at fixed points and the judge's model.

a ledger; an open-question map

In the Navier-Stokes work, a list of a manuscript's moves and their dependencies, and a list of its open questions.

a tier

A size class of model, small open or frontier.