Conjecture run panel:
Which open hypothesis engine should a field trust?
A growing number of open systems claim to generate scientific hypotheses: seven are surveyed here. None has been compared with another on the same literature, with the same prompt, under the same rules, and almost all of them mark their own homework. This is a harness that makes them comparable, points them at a corpus you supply, and ends every run with verification performed by something that did not do the generating.
Rationale
Same literature, same question, same rules
The published record on these systems cannot be compared, because nothing in it is held constant: each engine is demonstrated on its authors’ choice of corpus, with its authors’ prompt, through its own retrieval path, and judged, usually, by the model that did the generating. Three absences, specifically:
| No common ground | A different corpus and prompt per system, so a published result cannot say whether an architecture is better or merely better-suited to the example its authors chose. |
|---|---|
| Self-verification | Reflection loops, simulated debate, self-scoring: the model that generated is the model that judges. |
| No ground truth | Quality is assessed by asking a model whether it is impressed. Where a field holds verified evidence labels, a verifier can be scored instead of admired. |
The harness removes every difference except the one worth measuring. Engines read the literature only through one read-only corpus API, with no web access and no private retrieval, answer the identical brief, and return one output schema. What may differ is architecture and model family. What may not differ is the question, the evidence, or the shape of the answer. And the verdict on a hypothesis is never given by the family that wrote it, because we measured what self-judging costs:
Pipeline
How a run works
1 CORPUS seed bibliography → citation expansion → clustering → full-text mining
you supply the seed; the harness builds the universe and the search index
2 SEEDS open questions the field states about itself, plus gaps in its claim registry
each becomes a structured brief every engine receives identically
3 GENERATION N engines, one shared prompt, one read-only corpus API, no web access
engines retrieve agentically; retrieval behaviour is part of what is compared
4 MERGE cluster across engines, keep provenance
independent arrival at the same claim is signal; contradictions are gold
5 VERIFICATION decompose each claim into steps, audit every step against the cited papers,
certify or decline, by a family that did not generate
6 OUTPUT experiment cards, a scoreboard, and an audit database with the queries
that regenerate every published number
Getting started
What you need to start
One thing: a seed bibliography for the field you care about. Everything else is built from it. The Corpus page walks the full method, stage by stage, and takes the upload.
Build the corpus for a run
An engine is only as good as what it is allowed to read. This is the method that built the biophoton pack, reproduced end to end: a curated seed bibliography becomes a mapped publication universe, a clustered field structure, a mined statement corpus and a searchable knowledgebase. Bring the seed; the harness does the rest.
Intake
1. Intake: your seed bibliography
Upload
A Zotero or Mendeley export, a reference-manager CSV, a BibTeX or RIS file, or a plain list of DOIs. Parsed in your browser: nothing is uploaded anywhere.
Run parameters
Closed-access and hand-collected PDFs
Stage I harvests open-access PDFs automatically. Everything else (paywalled papers, book chapters, scans, anything obtained by hand) has to be supplied. Drop them here to generate the manifest and the ingest commands. Files are read for name and size only; nothing is uploaded.
Method
2. The method, stage by stage
Each stage is idempotent and cached, so a rerun is free and a failure resumes rather than
restarts. The hub's corpus builder (builder/run.py) runs A, B, C, E, the index, I, J
and K; clustering (D) and the claim registry (M) are a pack's own tools, since engines never read
the first and the second is curation.
| Stage | What it does | Output |
|---|---|---|
| A resolve | Match every seed reference to a canonical work id. Unresolvable entries are reported, not silently dropped. | seed work ids |
| B expand | Backward and forward citation expansion to the chosen hop depth, with pruning. This is what turns a reading list into a field. | the publication universe |
| C normalise | Works, authors, institutions, topics and citation edges into one relational store. | fieldmap.sqlite |
| D network pack tooling | Co-authorship, bibliographic-coupling and co-citation graphs, then Leiden community detection: chosen over Louvain because Leiden guarantees well-connected communities, which Louvain does not (Traag, Waltman & van Eck 2019, Sci. Rep. 9:5233, 10.1038/s41598-019-41695-z). Cluster membership is a function of shared references and nothing else: no keyword filter, no model, no editorial hand. | communities, sub-strands |
| E openness optional | Per-work and per-author open-access and open-data scoring. An analysis overlay, not a prerequisite: no engine reads it and the corpus API does not expose it. Run it if you want to characterise how open the field is; skip it and everything downstream still works. | openness overlay |
| I harvest | Fetch every open-access PDF the universe contains. Resumable. | local full texts |
| J mine | Extract text and mine the sentences where the field states what it does not know: the raw material for run seeds. | statement corpus |
| K knowledgebase | One joined, FTS5-searchable store over titles, abstracts and full text. This is what the corpus API serves to engines. | knowledgebase.sqlite |
| M registry pack curation | The field's canonical claims as scoped hypotheses with explicit nulls, linked to every sentence that supports, refutes or discusses them. Optional, but it is what makes ground-truth scoring possible. | claim registry + evidence |
Interface
3. What the corpus API exposes
Engines never touch the database. They see four calls, and nothing else.
search(query, limit, year_max) full-text over title, abstract and body
get_work(work_id) one record
statements(work_id, kind) mined open-problem sentences
neighbors(work_id) citing / cited within the universe
year_max exists so novelty can be scored honestly: restrict an engine to
literature before a cutoff, then check its "new" hypotheses against what was actually published
afterwards.
Sources & licences
Everything the harness reads or runs, and the terms it comes under. One engine carries a licence that restricts what may be published from its output; it is flagged.
Data
Corpus and data
| Source | What it provides | Licence / terms |
|---|---|---|
| OpenAlex | The publication universe and citation graph. In the biophoton pack: 18,355 works, 39,312 authors, ~265,000 intra-universe edges. | CC0 |
| Your seed bibliography | The starting point for expansion. Metadata only. | Yours |
| Full texts | Open-access PDFs, mined at sentence level. | Publisher © retained short quoted excerpts, attributed |
| Derived field map | Clustering and openness overlays for the biophoton pack. | CC0 1.0 10.5281/zenodo.21466492 |
| Claim registry | Scoped claims, verified stance labels, entailment tier: the ground truth. | CC0 |
Engines
Hypothesis engines
Licence, status and, behind Learn more on each row, what the architecture does, which seam this harness replaces to point it at our corpus, and the sources the description draws on.
Codex CLISingle agent + tools Vendor CLILearn more
The control condition, and deliberately the least clever thing here: one agent, one pass, tool access to the corpus. It searches, reads, searches again, writes. No tournament, no critic, no second opinion.
It exists so every multi-agent architecture has something to beat. An elaborate system that cannot outperform one agent with the same tools has not shown that its elaboration does anything. Running the identical architecture on two model families also separates "this engine is good" from "this model is good": which turned out to matter more than expected.
- OpenAI Codex CLI (
codex): vendor documentation. - Implementation:
engines/codex_solo/(manifest and adapter).
Claude subagentsSingle agent + tools Vendor CLILearn more
Same architecture as the Codex engine, other model family: that is the entire point of running both. Any difference in their output is a model-family difference, not a scaffolding one.
One honest limitation: it runs through agents dispatched inside a Claude Code session, because a
nested claude -p cannot authenticate from inside one. The hypotheses are genuine
inference; the engine is not reproducible by someone cloning the repo without that session. An
Anthropic API key turns it into an ordinary scripted engine.
- Anthropic Claude Code CLI: vendor documentation.
- Implementation:
engines/claude_agent/; status detail inconnectors/STATUS.md.
Codex over MCPSingle agent + MCP tools Vendor CLILearn more
The Codex engine with one change: it reads the corpus through the hub’s MCP server instead of a shell command. The pair isolates the retrieval route, and this adapter is the template for plugging in any agent that speaks the Model Context Protocol.
mcp_server.py serves the four corpus calls as tools.- Implementation:
engines/codex_mcp/.
Reference adapterNo model HubLearn more
Proposes that the best-matching findings for a question replicate. It exists so the contract can be exercised end to end with no model and no key, and as a floor: any engine worth running should beat it on every judged axis.
- Implementation:
engines/reference/.
LLNL Open AI Co-ScientistElo debate tournament MITLearn more
An open reimplementation of Google's AI co-scientist. Six agents under a supervisor: Generation proposes, Reflection reviews, Ranking runs a pairwise tournament where two hypotheses are argued against each other and the winner takes Elo points, Evolution breeds new hypotheses from survivors, Proximity maps similarity, Meta-Review summarises. Hypotheses carry an Elo score (from 1200) and a parent lineage, so you can trace which ancestors a surviving idea came from.
The interesting claim is that argument filters better than scoring: rather than asking a model to rate a hypothesis, make two compete and keep the winner.
ArxivSearchTool.search_papers → our corpus.
As shipped its arXiv results never reach the agents, so the connector must also wire retrieval into
the agent prompts: swapping the search alone would change nothing.- LLNL, Open AI Co-Scientist: Hypothesis Evolution System, repository README (open-source implementation of Google's AI Co-Scientist).
- Architecture read from
app/agents.pyandapp/tools/arxiv_search.pyin the cloned source.
HypoGeniCScored on labelled data MITLearn more
The only engine whose hypotheses are checked against outcomes rather than opinions. It keeps a bank of natural-language hypotheses and updates it like a bandit problem: each is applied to labelled training examples, earns a reward for how well it predicts them, and the bank keeps what performs. A companion mode (HypoRefine) folds literature into the same loop.
The constraint is the point: it needs a supervised dataset. Powerful where labels exist, inapplicable where they do not: which is most open questions in a field like this one.
LiteratureAgent(paper_infos=[{title, summary}])
→ our corpus records, bypassing GROBID and the PDF pipeline entirely.- Zhou et al., Hypothesis Generation with Large Language Models:
arxiv.org/abs/2404.04326 - HypoRefine, literature-integrated generation:
arxiv.org/abs/2410.17309 - HypoBench:
arxiv.org/abs/2504.11524
FutureHouse RobinAssays, Bradley-Terry Apache-2.0Learn more
Built for experimental biology, and it shows in the output shape: rather than proposing a claim it proposes an assay, the measurement you would actually run, then ranks candidates by pairwise comparison aggregated with a Bradley-Terry model. The published demonstration ran the loop end to end in drug discovery, including analysis of real experimental data.
That framing is the closest of any engine to the experiment cards this harness produces, which is why it is prioritised despite needing a fork.
call_platform() in robin/utils.py: one function, called four times, reaching the proprietary Edison platform. Replaced with a
corpus-backed implementation returning the same dict shape, which removes the paid dependency for
the hypothesis stage.- FutureHouse, Robin: A multi-agent system for automating scientific discovery:
arxiv.org/abs/2505.13400 - Seam identified in the cloned source at
robin/robin/utils.py.
SciAgentsKnowledge-graph paths Apache-2.0Learn more
The one engine whose native input is a graph rather than a search box. It builds an ontological knowledge graph, samples a path between two concepts, and has a team (Ontologist, two Scientists, a Critic) expand that path into a structured proposal with hypothesis, mechanism, design principles and a novelty argument. The premise is that novelty lives in connections a literature has not yet made, and sampling paths surfaces them.
Conceptually the best fit here, since a citation-coupling graph is exactly what we already have.
Practically the most work: it wants a .graphml and a node-embedding pickle, not a
search tool, and the upstream code has not changed materially in some time.
.graphml plus matching node embeddings.- Ghafarollahi & Buehler, SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning.
- Graph construction: Buehler et al., 2024,
10.1088/2632-2153/ad7228
Sakana AI Scientist v2Reflection loops, end to end not open sourceLearn more
The most ambitious of the set. It takes a topic and runs the whole cycle: ideation, experiments, analysis, written manuscript. Only its ideation stage is used here. That stage generates candidate ideas then iteratively reflects on them, rewriting each across several passes before finalising. v2 drops v1's reliance on human-authored templates and uses an agentic tree search over the experiment space.
Its ideation stage is unusually easy to isolate (one self-contained file with its own entry point), which is why it is Tier 1 despite the licence problem.
SemanticScholarSearchTool → a
BaseTool subclass over our corpus, swapped into the module-level tool list. Verified by
isinstance against their abstract base class.- Sakana AI, The AI Scientist-v2: repository README and paper (
pub.sakana.ai/ai-scientist-v2/paper). - Tool contract from
ai_scientist/tools/base_tool.py; licence text fromLICENSEin the cloned source.
OpenScientist (LBNL)Generate-and-test platform Apache-2.0Learn more
A platform rather than a library: give it data files and a research question and it runs an iterative generate-and-test loop, recording hypotheses, findings and consulted literature in a relational store. It is the only system here with genuine relational provenance (findings link to the hypotheses they bear on and to the papers behind them), which is the property this harness values most and which most engines lack.
The cost is that it is the least separable: no standalone hypothesis-generation entry point, so integration means running the platform and ingesting our corpus into its own literature store.
- Reese et al. (Lawrence Berkeley National Laboratory), OpenScientist:
openscientist.io. - Preprint: medRxiv
10.64898/2026.03.15.26348338.
Method
Verification
| Source | What is taken from it | Licence |
|---|---|---|
| theoria | The certify-or-decline discipline: stepwise proof, independent judges per step, certify only if every load-bearing step survives, and an audit database with the queries that regenerate every number. Wired in and running: see the Verification page. | Apache-2.0 |
Configure a run
Pick the corpus, the engines that generate and the families that verify. The harness emits the commands. Engines needing an API key are shown but not selectable until one exists.
Corpus
Depth and cost
How hard each engine works. Depth scales hypotheses per seed, reasoning effort, verification, and therefore cost.
Engines: generation
Estimated cost is for the seeds selected below, at the chosen depth.
Verification: judge families
Seeds
API keys
Stored in
hub/.env on this machine: file mode 0600, gitignored, shown back only as the
last four characters. A key unlocks its engines for future runs; the runners for the keyed
engines are still being wired, so saving a key changes their status but does not yet make the
button run them.
Trigger
What the button can genuinely run unattended today. Anything else stays in the command list below. theoria is scoped, never swept: it runs only the claims ticked in its list, and both this button and the adapter itself refuse to start on an empty or implicit selection.
Estimated cost
Provider prices (edit to match your plan)
Commands
…This run would cover
Hypotheses:
questions the field states about itself, every engine that has run on them, and every verdict against what they produced.
The ledger
Every hypothesis, and every verdict against it
One row per hypothesis. The judge columns are our own audit: does each reasoning step hold against the corpus it cites. The theoria column is independent of all of it: theoria never sees the citations or which engine wrote the hypothesis, only the arithmetic, which it re-derives from scratch. A hypothesis can pass our audit and still fail theoria's: where the two disagree is the reason for running both, and the ledger flags every such case as it lands.
Output
Generation
Signal
Cross-engine convergence
Where two independent model families reached the same claim in different vocabulary. Independent arrival is weak evidence that a hypothesis is findable in the literature rather than invented.
Signal
Conflicts: discriminating experiments
Pairs asserting incompatible things about the same measurand. One being right makes the other wrong, so each pair defines an experiment that would settle it.
Output
Certified hypotheses
Survived step-by-step audit under unanimity across all judge families, and judged not to be restatements of the existing registry.
Output
Declined
Published with reasons: the declines are the point of the exercise.
Verification
Every run ends here, and never with the model that generated. The discipline is adapted from theoria; the verifiers are then themselves measured, which is possible wherever a pack holds labelled ground truth, which no generator project has.
Discipline
1. The certification discipline
theoria is a verified-reasoning harness, not a benchmark suite. A claim must be emitted as a stepwise proof; independent judges audit every step; the answer is certified or declined, never quietly accepted. We adopt that discipline with one addition its author did not need: a citation step must resolve to a work that exists in this corpus and actually assert what the step claims.
| Stage | What it checks | Cost |
|---|---|---|
| Mechanical | Citations exist; scope complete; experiment specific enough to cost; registry restatement flagged. | Free: declines a class before any judge is paid |
| Judged | Every rationale step, with the cited abstracts inlined. Does the cited work actually say this? Would the experiment discriminate the claim from its null? | One pass per judge family |
| Verdict | certified · pedantic (sound, non-load-bearing flaw) · declined (failing step named) | Unanimity required |
Independent verifier
2. theoria, actually run
Sections 1 and 3 to 7 describe our own judges. This section is different: it is theoria itself, cloned and executed (its solver, interpreter, formalizer, per-step judges, pedantry filter and convention lift), not our imitation of its discipline.
It is given something deliberately narrow. We extract the steps an engine itself marked as computation, plus the estimand, and pose them as one self-contained proposition. theoria never sees the citations, the seed question, or which engine wrote it. It re-derives every number from scratch and rejects any step resting on an unstated premise, then answers CORRECT or INCORRECT with the specific error.
That makes it independent of our judges in the way that matters. Our judges ask a question about this corpus: does the cited paper support the step citing it. theoria asks whether the reasoning is valid at all. A hypothesis can be scrupulously cited and still arithmetically wrong, which is exactly the failure class the run turned up on the physics side.
What theoria can answer, and what each outcome means
theoria separates two questions that are usually conflated: did I produce a sound proof (its own audit of itself) and what does that proof conclude (the answer). Both are in every run record, and both are needed to read a verdict correctly.
| Field | Values | What it means |
|---|---|---|
answer | free text: for our claims, CORRECT / INCORRECT + the specific error | The conclusion of the proof: what theoria established about the claim it was posed. Extracted from the final proof state. |
verified | JUDGE-PASSED / REJECTED | Whether theoria's own proof survived its per-step audit: one judge per step, the judge's kind chosen by the step's justification type. REJECTED means theoria could not establish anything either way: not that the claim is false. |
verified_unconditionally | true / false | True only if the proof passed with zero added assumptions. False-but-verified means the convention-lift judge had to grant something (next row). |
verified_under_assumptions | list of named conventions | Rejected steps that were legitimate but rescued by naming an explicit, sourced convention (e.g. a standard sign convention). Each is recorded with its step and source, so the assumption is auditable rather than silent. |
per-step verdicts | accepted / rejected + reason, per step | The audit itself. A citation step's judge checks the invoked result exists and is applied correctly; a computation step's judge re-derives the number; a problem_given judge checks the premise really is stated in the problem. |
pedantry | is_pedantic + reason | A screen over failed verdicts: a rejection that is a technicality (formatting, an unstated-but-universal convention) is overturned here rather than sinking the proof. Substantive rejections pass through. |
correct | true / false | A naive substring match against an expected answer: upstream documents it as for scanning only. Meaningless for our custom claims (there is no expected string) and never used here. |
How those become the ledger's theoria column
| Ledger verdict | Condition | Reading |
|---|---|---|
| certified | proof verified and answer CORRECT | theoria produced an audited proof that the hypothesis's arithmetic holds. |
| declined | proof verified and answer INCORRECT | An audited proof that the arithmetic is wrong, with the specific error. A soundly-proved INCORRECT is a decline: conflating verified with approval would invert exactly these cases. |
| inconclusive | proof not verified, or no readable answer | theoria could not establish the claim either way. Not evidence against the hypothesis. |
| error | timeout or crash | The run did not complete. Re-queued, not counted. |
| queued | claim extracted, run pending | Waiting its turn: each claim takes tens of minutes. |
| no arithmetic | no computation step in the hypothesis | A purely empirical proposal: nothing for theoria to re-derive. Blank would read as an unrun check, so it is stated. |
Fidelity to the upstream setup: stated, not assumed
| Audited theoria setup | This machine | Consequence |
|---|---|---|
| Codex on every role, Claude formalizer | Codex (gpt-5.5) on every role incl. formalizer: a nested Claude CLI cannot authenticate inside the session that drives this harness | Single-family verification. Recorded because self-preference is this project's own headline measurement. |
| Each problem in a Docker sandbox; agents told about the container | --no-docker (Docker not installed); agents use the Codex CLI's native read-only sandbox on the host | Isolation is weaker (read-only sandbox rather than a container). Tool execution works: verified by probe. |
| Environment preamble matches the sandbox | Pass 1 did not: upstream's Debian-container description was handed to agents on a bare host. Measured effect: 52 calls, zero tool invocations: arithmetic "re-derived" by unexecuted reasoning. Pass 1 is archived, not reported. Pass 2 runs with a preamble describing the real environment and instructing executed python3 checks. | The pass-1 verdicts may still be right, but they were not produced the way the harness claims, so they do not count. |
Consumption, measured. Per claim: 31 to 92 LLM calls, mean 6.3M input tokens
(4.2M of it prompt-cache reads) and 165k output, 22 to 35 minutes wall-clock. All 50 eligible
claims project to roughly 315M input / 8M output tokens. Billed through the Codex CLI
subscription: the run records' cost fields read zero throughout. Because a full sweep is
about a day of continuous compute, runs are now launched per named claim: the adapter
refuses --run without --only C001,… or an explicit
--all, and the Run page's trigger submits only ticked claims.
Benchmark
3. Verifier benchmark, scored against known labels
Where a pack holds a verified claim register, the verifiers themselves can be scored: blinded items, each a scoped claim plus one verbatim sentence whose stance was already adversarially verified, with refutations over-represented so that a verifier that simply agrees with the field scores badly. Sentences quoted publicly are excluded, so nobody can answer from memory.
Headline