Conjecture hypothesis hub · run panel

Conjecture run panel:

Which open hypothesis engine should a field trust?

A growing number of open systems claim to generate scientific hypotheses: seven are surveyed here. None has been compared with another on the same literature, with the same prompt, under the same rules, and almost all of them mark their own homework. This is a harness that makes them comparable, points them at a corpus you supply, and ends every run with verification performed by something that did not do the generating.

CompareAny open engine, one interface, one prompt, one output format
Any corpusBring a seed bibliography; the harness builds the literature universe around it
VerifyIndependent certify-or-decline audit, never by the generating model
MeasureScored against ground truth where a field has any

Rationale

Same literature, same question, same rules

The published record on these systems cannot be compared, because nothing in it is held constant: each engine is demonstrated on its authors’ choice of corpus, with its authors’ prompt, through its own retrieval path, and judged, usually, by the model that did the generating. Three absences, specifically:

No common groundA different corpus and prompt per system, so a published result cannot say whether an architecture is better or merely better-suited to the example its authors chose.
Self-verificationReflection loops, simulated debate, self-scoring: the model that generated is the model that judges.
No ground truthQuality is assessed by asking a model whether it is impressed. Where a field holds verified evidence labels, a verifier can be scored instead of admired.

The harness removes every difference except the one worth measuring. Engines read the literature only through one read-only corpus API, with no web access and no private retrieval, answer the identical brief, and return one output schema. What may differ is architecture and model family. What may not differ is the question, the evidence, or the shape of the answer. And the verdict on a hypothesis is never given by the family that wrote it, because we measured what self-judging costs:

The finding that forced the design. Both judge families in the first pack’s first run (biophotons) preferred hypotheses written by their own family: the Claude judges by 12 points, the Codex judge by 40 (45% pass rate on its own family’s output against 5% on the other’s). If one engine simply wrote better hypotheses, both judges would have ranked it higher; instead each ranked its own higher, which is the signature of self-preference rather than quality. A self-verified certification rate is therefore not a measurement, and single-family pipelines’ published scores are not comparable with each other. See the measurement →

Pipeline

How a run works

1  CORPUS        seed bibliography  →  citation expansion  →  clustering  →  full-text mining
                 you supply the seed; the harness builds the universe and the search index

2  SEEDS         open questions the field states about itself, plus gaps in its claim registry
                 each becomes a structured brief every engine receives identically

3  GENERATION    N engines, one shared prompt, one read-only corpus API, no web access
                 engines retrieve agentically; retrieval behaviour is part of what is compared

4  MERGE         cluster across engines, keep provenance
                 independent arrival at the same claim is signal; contradictions are gold

5  VERIFICATION  decompose each claim into steps, audit every step against the cited papers,
                 certify or decline, by a family that did not generate

6  OUTPUT        experiment cards, a scoreboard, and an audit database with the queries
                 that regenerate every published number

Getting started

What you need to start

One thing: a seed bibliography for the field you care about. Everything else is built from it. The Corpus page walks the full method, stage by stage, and takes the upload.

This panel is serving the pack . The hub is field-agnostic; a pack supplies the corpus, the questions and the run history. The first pack, biophotons (ultraweak photon emission), was chosen because the field is small enough to map completely, contested enough that hypothesis quality matters, and already carries a verified claim registry that can serve as ground truth. The Results and Verification pages report whichever pack is loaded.

Build the corpus for a run

An engine is only as good as what it is allowed to read. This is the method that built the biophoton pack, reproduced end to end: a curated seed bibliography becomes a mapped publication universe, a clustered field structure, a mined statement corpus and a searchable knowledgebase. Bring the seed; the harness does the rest.

Intake

1. Intake: your seed bibliography

Upload

A Zotero or Mendeley export, a reference-manager CSV, a BibTeX or RIS file, or a plain list of DOIs. Parsed in your browser: nothing is uploaded anywhere.

Run parameters

Closed-access and hand-collected PDFs

Stage I harvests open-access PDFs automatically. Everything else (paywalled papers, book chapters, scans, anything obtained by hand) has to be supplied. Drop them here to generate the manifest and the ingest commands. Files are read for name and size only; nothing is uploaded.

Method

2. The method, stage by stage

Each stage is idempotent and cached, so a rerun is free and a failure resumes rather than restarts. The hub's corpus builder (builder/run.py) runs A, B, C, E, the index, I, J and K; clustering (D) and the claim registry (M) are a pack's own tools, since engines never read the first and the second is curation.

StageWhat it doesOutput
A resolveMatch every seed reference to a canonical work id. Unresolvable entries are reported, not silently dropped.seed work ids
B expandBackward and forward citation expansion to the chosen hop depth, with pruning. This is what turns a reading list into a field.the publication universe
C normaliseWorks, authors, institutions, topics and citation edges into one relational store.fieldmap.sqlite
D network pack toolingCo-authorship, bibliographic-coupling and co-citation graphs, then Leiden community detection: chosen over Louvain because Leiden guarantees well-connected communities, which Louvain does not (Traag, Waltman & van Eck 2019, Sci. Rep. 9:5233, 10.1038/s41598-019-41695-z). Cluster membership is a function of shared references and nothing else: no keyword filter, no model, no editorial hand.communities, sub-strands
E openness optionalPer-work and per-author open-access and open-data scoring. An analysis overlay, not a prerequisite: no engine reads it and the corpus API does not expose it. Run it if you want to characterise how open the field is; skip it and everything downstream still works.openness overlay
I harvestFetch every open-access PDF the universe contains. Resumable.local full texts
J mineExtract text and mine the sentences where the field states what it does not know: the raw material for run seeds.statement corpus
K knowledgebaseOne joined, FTS5-searchable store over titles, abstracts and full text. This is what the corpus API serves to engines.knowledgebase.sqlite
M registry pack curationThe field's canonical claims as scoped hypotheses with explicit nulls, linked to every sentence that supports, refutes or discusses them. Optional, but it is what makes ground-truth scoring possible.claim registry + evidence
The registry is the expensive part, and the valuable one. In the biophoton pack it holds 15 scoped claims, 861 adversarially verified stance labels and 467 entailment judgments. Without it you can still generate and certify hypotheses; with it you can also measure whether an engine rediscovers what the field already believes, and whether a verifier is any good. Most fields have no such artefact, which is why verifier benchmarking has not been done before.

Interface

3. What the corpus API exposes

Engines never touch the database. They see four calls, and nothing else.

search(query, limit, year_max)   full-text over title, abstract and body
get_work(work_id)                one record
statements(work_id, kind)        mined open-problem sentences
neighbors(work_id)               citing / cited within the universe

year_max exists so novelty can be scored honestly: restrict an engine to literature before a cutoff, then check its "new" hypotheses against what was actually published afterwards.

Sources & licences

Everything the harness reads or runs, and the terms it comes under. One engine carries a licence that restricts what may be published from its output; it is flagged.

Data

Corpus and data

SourceWhat it providesLicence / terms
OpenAlexThe publication universe and citation graph. In the biophoton pack: 18,355 works, 39,312 authors, ~265,000 intra-universe edges.CC0
Your seed bibliographyThe starting point for expansion. Metadata only.Yours
Full textsOpen-access PDFs, mined at sentence level.Publisher © retained short quoted excerpts, attributed
Derived field mapClustering and openness overlays for the biophoton pack.CC0 1.0 10.5281/zenodo.21466492
Claim registryScoped claims, verified stance labels, entailment tier: the ground truth.CC0

Engines

Hypothesis engines

Licence, status and, behind Learn more on each row, what the architecture does, which seam this harness replaces to point it at our corpus, and the sources the description draws on.

EngineArchitectureLicence Status
Codex CLISingle agent + tools Vendor CLILearn more

The control condition, and deliberately the least clever thing here: one agent, one pass, tool access to the corpus. It searches, reads, searches again, writes. No tournament, no critic, no second opinion.

It exists so every multi-agent architecture has something to beat. An elaborate system that cannot outperform one agent with the same tools has not shown that its elaboration does anything. Running the identical architecture on two model families also separates "this engine is good" from "this model is good": which turned out to matter more than expected.

Seam: none: calls our corpus API natively.
Sources
  • OpenAI Codex CLI (codex): vendor documentation.
  • Implementation: engines/codex_solo/ (manifest and adapter).
Claude subagentsSingle agent + tools Vendor CLILearn more

Same architecture as the Codex engine, other model family: that is the entire point of running both. Any difference in their output is a model-family difference, not a scaffolding one.

One honest limitation: it runs through agents dispatched inside a Claude Code session, because a nested claude -p cannot authenticate from inside one. The hypotheses are genuine inference; the engine is not reproducible by someone cloning the repo without that session. An Anthropic API key turns it into an ordinary scripted engine.

Seam: none: calls our corpus API natively.
Sources
  • Anthropic Claude Code CLI: vendor documentation.
  • Implementation: engines/claude_agent/; status detail in connectors/STATUS.md.
Codex over MCPSingle agent + MCP tools Vendor CLILearn more

The Codex engine with one change: it reads the corpus through the hub’s MCP server instead of a shell command. The pair isolates the retrieval route, and this adapter is the template for plugging in any agent that speaks the Model Context Protocol.

Seam: none: mcp_server.py serves the four corpus calls as tools.
Sources
  • Implementation: engines/codex_mcp/.
Reference adapterNo model HubLearn more

Proposes that the best-matching findings for a question replicate. It exists so the contract can be exercised end to end with no model and no key, and as a floor: any engine worth running should beat it on every judged axis.

Sources
  • Implementation: engines/reference/.
LLNL Open AI Co-ScientistElo debate tournament MITLearn more

An open reimplementation of Google's AI co-scientist. Six agents under a supervisor: Generation proposes, Reflection reviews, Ranking runs a pairwise tournament where two hypotheses are argued against each other and the winner takes Elo points, Evolution breeds new hypotheses from survivors, Proximity maps similarity, Meta-Review summarises. Hypotheses carry an Elo score (from 1200) and a parent lineage, so you can trace which ancestors a surviving idea came from.

The interesting claim is that argument filters better than scoring: rather than asking a model to rate a hypothesis, make two compete and keep the winner.

Seam: ArxivSearchTool.search_papers → our corpus. As shipped its arXiv results never reach the agents, so the connector must also wire retrieval into the agent prompts: swapping the search alone would change nothing.
Sources
  • LLNL, Open AI Co-Scientist: Hypothesis Evolution System, repository README (open-source implementation of Google's AI Co-Scientist).
  • Architecture read from app/agents.py and app/tools/arxiv_search.py in the cloned source.
HypoGeniCScored on labelled data MITLearn more

The only engine whose hypotheses are checked against outcomes rather than opinions. It keeps a bank of natural-language hypotheses and updates it like a bandit problem: each is applied to labelled training examples, earns a reward for how well it predicts them, and the bank keeps what performs. A companion mode (HypoRefine) folds literature into the same loop.

The constraint is the point: it needs a supervised dataset. Powerful where labels exist, inapplicable where they do not: which is most open questions in a field like this one.

Seam: LiteratureAgent(paper_infos=[{title, summary}]) → our corpus records, bypassing GROBID and the PDF pipeline entirely.
Sources
  • Zhou et al., Hypothesis Generation with Large Language Models: arxiv.org/abs/2404.04326
  • HypoRefine, literature-integrated generation: arxiv.org/abs/2410.17309
  • HypoBench: arxiv.org/abs/2504.11524
FutureHouse RobinAssays, Bradley-Terry Apache-2.0Learn more

Built for experimental biology, and it shows in the output shape: rather than proposing a claim it proposes an assay, the measurement you would actually run, then ranks candidates by pairwise comparison aggregated with a Bradley-Terry model. The published demonstration ran the loop end to end in drug discovery, including analysis of real experimental data.

That framing is the closest of any engine to the experiment cards this harness produces, which is why it is prioritised despite needing a fork.

Seam: call_platform() in robin/utils.py: one function, called four times, reaching the proprietary Edison platform. Replaced with a corpus-backed implementation returning the same dict shape, which removes the paid dependency for the hypothesis stage.
Sources
  • FutureHouse, Robin: A multi-agent system for automating scientific discovery: arxiv.org/abs/2505.13400
  • Seam identified in the cloned source at robin/robin/utils.py.
SciAgentsKnowledge-graph paths Apache-2.0Learn more

The one engine whose native input is a graph rather than a search box. It builds an ontological knowledge graph, samples a path between two concepts, and has a team (Ontologist, two Scientists, a Critic) expand that path into a structured proposal with hypothesis, mechanism, design principles and a novelty argument. The premise is that novelty lives in connections a literature has not yet made, and sampling paths surfaces them.

Conceptually the best fit here, since a citation-coupling graph is exactly what we already have. Practically the most work: it wants a .graphml and a node-embedding pickle, not a search tool, and the upstream code has not changed materially in some time.

Seam: none written: requires exporting our concept graph as .graphml plus matching node embeddings.
Sources
  • Ghafarollahi & Buehler, SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning.
  • Graph construction: Buehler et al., 2024, 10.1088/2632-2153/ad7228
Sakana AI Scientist v2Reflection loops, end to end not open sourceLearn more

The most ambitious of the set. It takes a topic and runs the whole cycle: ideation, experiments, analysis, written manuscript. Only its ideation stage is used here. That stage generates candidate ideas then iteratively reflects on them, rewriting each across several passes before finalising. v2 drops v1's reliance on human-authored templates and uses an agentic tree search over the experiment space.

Its ideation stage is unusually easy to isolate (one self-contained file with its own entry point), which is why it is Tier 1 despite the licence problem.

Seam: SemanticScholarSearchTool → a BaseTool subclass over our corpus, swapped into the module-level tool list. Verified by isinstance against their abstract base class.
Licence gate. This is the one engine here that is not open source. It ships under "The AI Scientist Source Code License", derived from the Responsible AI licence family. Internal benchmarking is permitted, but any generated manuscript or technical report must prominently disclose machine generation, and the restrictions propagate into downstream agreements. Its output must be labelled and legally reviewed before publication: treat anything it produces as encumbered until that review happens. Every other engine here is MIT or Apache-2.0.
Sources
  • Sakana AI, The AI Scientist-v2: repository README and paper (pub.sakana.ai/ai-scientist-v2/paper).
  • Tool contract from ai_scientist/tools/base_tool.py; licence text from LICENSE in the cloned source.
OpenScientist (LBNL)Generate-and-test platform Apache-2.0Learn more

A platform rather than a library: give it data files and a research question and it runs an iterative generate-and-test loop, recording hypotheses, findings and consulted literature in a relational store. It is the only system here with genuine relational provenance (findings link to the hypotheses they bear on and to the papers behind them), which is the property this harness values most and which most engines lack.

The cost is that it is the least separable: no standalone hypothesis-generation entry point, so integration means running the platform and ingesting our corpus into its own literature store.

Seam: none written: the airgapped route ingests our works into its local literature mirror rather than querying PubMed.
Sources
  • Reese et al. (Lawrence Berkeley National Laboratory), OpenScientist: openscientist.io.
  • Preprint: medRxiv 10.64898/2026.03.15.26348338.

Method

Verification

SourceWhat is taken from itLicence
theoriaThe certify-or-decline discipline: stepwise proof, independent judges per step, certify only if every load-bearing step survives, and an audit database with the queries that regenerate every number. Wired in and running: see the Verification page.Apache-2.0

Configure a run

Pick the corpus, the engines that generate and the families that verify. The harness emits the commands. Engines needing an API key are shown but not selectable until one exists.

Corpus

Depth and cost

How hard each engine works. Depth scales hypotheses per seed, reasoning effort, verification, and therefore cost.

Engines: generation

Estimated cost is for the seeds selected below, at the chosen depth.

Verification: judge families

A candidate is certified only if every selected family certifies it.

Seeds

API keys

Stored in hub/.env on this machine: file mode 0600, gitignored, shown back only as the last four characters. A key unlocks its engines for future runs; the runners for the keyed engines are still being wired, so saving a key changes their status but does not yet make the button run them.

Trigger

What the button can genuinely run unattended today. Anything else stays in the command list below. theoria is scoped, never swept: it runs only the claims ticked in its list, and both this button and the adapter itself refuse to start on an empty or implicit selection.

Estimated cost

Provider prices (edit to match your plan)

Commands

…

This run would cover

Hypotheses:

questions the field states about itself, every engine that has run on them, and every verdict against what they produced.

The ledger

Every hypothesis, and every verdict against it

One row per hypothesis. The judge columns are our own audit: does each reasoning step hold against the corpus it cites. The theoria column is independent of all of it: theoria never sees the citations or which engine wrote the hypothesis, only the arithmetic, which it re-derives from scratch. A hypothesis can pass our audit and still fail theoria's: where the two disagree is the reason for running both, and the ledger flags every such case as it lands.

Output

Generation

Signal

Cross-engine convergence

Where two independent model families reached the same claim in different vocabulary. Independent arrival is weak evidence that a hypothesis is findable in the literature rather than invented.

Signal

Conflicts: discriminating experiments

Pairs asserting incompatible things about the same measurand. One being right makes the other wrong, so each pair defines an experiment that would settle it.

Output

Certified hypotheses

Survived step-by-step audit under unanimity across all judge families, and judged not to be restatements of the existing registry.

Output

Declined

Published with reasons: the declines are the point of the exercise.

Verification

Every run ends here, and never with the model that generated. The discipline is adapted from theoria; the verifiers are then themselves measured, which is possible wherever a pack holds labelled ground truth, which no generator project has.

Discipline

1. The certification discipline

theoria is a verified-reasoning harness, not a benchmark suite. A claim must be emitted as a stepwise proof; independent judges audit every step; the answer is certified or declined, never quietly accepted. We adopt that discipline with one addition its author did not need: a citation step must resolve to a work that exists in this corpus and actually assert what the step claims.

StageWhat it checksCost
MechanicalCitations exist; scope complete; experiment specific enough to cost; registry restatement flagged.Free: declines a class before any judge is paid
JudgedEvery rationale step, with the cited abstracts inlined. Does the cited work actually say this? Would the experiment discriminate the claim from its null?One pass per judge family
Verdictcertified · pedantic (sound, non-load-bearing flaw) · declined (failing step named)Unanimity required

Independent verifier

2. theoria, actually run

Sections 1 and 3 to 7 describe our own judges. This section is different: it is theoria itself, cloned and executed (its solver, interpreter, formalizer, per-step judges, pedantry filter and convention lift), not our imitation of its discipline.

It is given something deliberately narrow. We extract the steps an engine itself marked as computation, plus the estimand, and pose them as one self-contained proposition. theoria never sees the citations, the seed question, or which engine wrote it. It re-derives every number from scratch and rejects any step resting on an unstated premise, then answers CORRECT or INCORRECT with the specific error.

That makes it independent of our judges in the way that matters. Our judges ask a question about this corpus: does the cited paper support the step citing it. theoria asks whether the reasoning is valid at all. A hypothesis can be scrupulously cited and still arithmetically wrong, which is exactly the failure class the run turned up on the physics side.

What theoria can answer, and what each outcome means

theoria separates two questions that are usually conflated: did I produce a sound proof (its own audit of itself) and what does that proof conclude (the answer). Both are in every run record, and both are needed to read a verdict correctly.

FieldValuesWhat it means
answerfree text: for our claims, CORRECT / INCORRECT + the specific errorThe conclusion of the proof: what theoria established about the claim it was posed. Extracted from the final proof state.
verifiedJUDGE-PASSED / REJECTEDWhether theoria's own proof survived its per-step audit: one judge per step, the judge's kind chosen by the step's justification type. REJECTED means theoria could not establish anything either way: not that the claim is false.
verified_unconditionallytrue / falseTrue only if the proof passed with zero added assumptions. False-but-verified means the convention-lift judge had to grant something (next row).
verified_under_assumptionslist of named conventionsRejected steps that were legitimate but rescued by naming an explicit, sourced convention (e.g. a standard sign convention). Each is recorded with its step and source, so the assumption is auditable rather than silent.
per-step verdictsaccepted / rejected + reason, per stepThe audit itself. A citation step's judge checks the invoked result exists and is applied correctly; a computation step's judge re-derives the number; a problem_given judge checks the premise really is stated in the problem.
pedantryis_pedantic + reasonA screen over failed verdicts: a rejection that is a technicality (formatting, an unstated-but-universal convention) is overturned here rather than sinking the proof. Substantive rejections pass through.
correcttrue / falseA naive substring match against an expected answer: upstream documents it as for scanning only. Meaningless for our custom claims (there is no expected string) and never used here.

How those become the ledger's theoria column

Ledger verdictConditionReading
certifiedproof verified and answer CORRECTtheoria produced an audited proof that the hypothesis's arithmetic holds.
declinedproof verified and answer INCORRECTAn audited proof that the arithmetic is wrong, with the specific error. A soundly-proved INCORRECT is a decline: conflating verified with approval would invert exactly these cases.
inconclusiveproof not verified, or no readable answertheoria could not establish the claim either way. Not evidence against the hypothesis.
errortimeout or crashThe run did not complete. Re-queued, not counted.
queuedclaim extracted, run pendingWaiting its turn: each claim takes tens of minutes.
no arithmeticno computation step in the hypothesisA purely empirical proposal: nothing for theoria to re-derive. Blank would read as an unrun check, so it is stated.

Fidelity to the upstream setup: stated, not assumed

Audited theoria setupThis machineConsequence
Codex on every role, Claude formalizerCodex (gpt-5.5) on every role incl. formalizer: a nested Claude CLI cannot authenticate inside the session that drives this harnessSingle-family verification. Recorded because self-preference is this project's own headline measurement.
Each problem in a Docker sandbox; agents told about the container--no-docker (Docker not installed); agents use the Codex CLI's native read-only sandbox on the hostIsolation is weaker (read-only sandbox rather than a container). Tool execution works: verified by probe.
Environment preamble matches the sandboxPass 1 did not: upstream's Debian-container description was handed to agents on a bare host. Measured effect: 52 calls, zero tool invocations: arithmetic "re-derived" by unexecuted reasoning. Pass 1 is archived, not reported. Pass 2 runs with a preamble describing the real environment and instructing executed python3 checks.The pass-1 verdicts may still be right, but they were not produced the way the harness claims, so they do not count.

Consumption, measured. Per claim: 31 to 92 LLM calls, mean 6.3M input tokens (4.2M of it prompt-cache reads) and 165k output, 22 to 35 minutes wall-clock. All 50 eligible claims project to roughly 315M input / 8M output tokens. Billed through the Codex CLI subscription: the run records' cost fields read zero throughout. Because a full sweep is about a day of continuous compute, runs are now launched per named claim: the adapter refuses --run without --only C001,… or an explicit --all, and the Run page's trigger submits only ticked claims.

Reading its verdict correctly. theoria reports whether it verified its own proof, separately from what that proof concluded. A soundly verified proof of INCORRECT is a decline of the hypothesis, not a pass. Conflating the two inverts the result: it did so here, and four hypotheses were briefly reported as certified whose answer was INCORRECT. The verdicts on this page were re-derived from the saved run JSONs after the fix.

Benchmark

3. Verifier benchmark, scored against known labels

Where a pack holds a verified claim register, the verifiers themselves can be scored: blinded items, each a scoped claim plus one verbatim sentence whose stance was already adversarially verified, with refutations over-represented so that a verifier that simply agrees with the field scores badly. Sentences quoted publicly are excluded, so nobody can answer from memory.

Headline

4. Do judges prefer their own model family?

Why the harness measures this. If one engine simply wrote better hypotheses, every judge would rank it higher. If instead each judge ranks its own family higher, a certification rate measures who judged as much as what was written. Almost every engine in the landscape verifies with the model that generated (reflection loops, simulated debate, self-scoring, data-driven reward), so cross-family verification is wired into the harness as the minimum valid design rather than offered as an option.