Engines and judges
Engines
Each engine answers the same brief through the same corpus interface. Built adapters have passed conformance: on a fictional corpus that no model can know from outside, they return valid output, cite only works in the corpus and fetch nothing from elsewhere.
- Claude subagent, single agent + corpus CLI. The Claude counterpart of codex-solo, same architecture, so scoreboard differences are model-family differences. Runs as a subagent dispatched in a Claude Code session.
Corpus access: command line. Licence: hub (AGPL-3.0-or-later). Status: ran on biophotons, 42 hypotheses. - Codex, single agent + corpus CLI. One Codex agent, one pass, corpus through corpus_api.py in a read-only shell. The control every multi-agent engine must beat, on a model family that is not Claude.
Corpus access: command line. Licence: hub (AGPL-3.0-or-later). Status: ran on biophotons, 40 hypotheses. - Codex, single agent + corpus over MCP. codex-solo with the corpus served by mcp_server.py instead of a shell, so the pair isolates the effect of the access route. The template for plugging in any MCP-speaking agent.
Corpus access: MCP server. Licence: hub (AGPL-3.0-or-later). Status: needs a local requirement. - Reference adapter (no model). Proposes that the best-matching findings replicate. Exercises the whole contract (MCP retrieval, HOF, citation checks) with no model and no key; a floor every engine should beat.
Corpus access: MCP server. Licence: hub (AGPL-3.0-or-later). Status: ready, conformance passed. - Sakana AI Scientist v2 (ideation stage). Ideation with reflection loops.
Corpus access: connector. Licence: The AI Scientist Source Code License 1.0 (not OSI open source). Status: designed, not yet built (API key; licence gate on output). - LLNL Open AI Co-Scientist. Elo tournament of simulated scientific debates.
Corpus access: connector. Licence: MIT. Status: designed, not yet built (OpenRouter key. As shipped, its search results never reach its agents; the connector wires them into the prompts, and that part is unrun). - HypoGeniC (ChicagoHAI). Data-driven hypothesis bank scored on labelled accuracy.
Corpus access: connector. Licence: MIT. Status: designed, not yet built (API key). - OpenScientist (LBNL). Generate-and-test platform with relational provenance.
Corpus access: not yet designed. Licence: Apache-2.0. Status: designed, not yet built (a platform, not a library; needs the corpus ingested into its Postgres mirror). - FutureHouse Robin (fork). Assay-framed output with Bradley-Terry ranking.
Corpus access: connector. Licence: Apache-2.0. Status: designed, not yet built (API key for the local reasoning calls). - SciAgents (MIT Buehler lab). Knowledge-graph traversal; its native input matches the field map.
Corpus access: not yet designed. Licence: Apache-2.0. Status: designed, not yet built (needs a .graphml and node-embedding export of the pack, not a search swap).
Judges
A candidate is certified only if every judge family certifies it. Judges audit each reasoning step against the abstracts of the works it cites.
- Claude judges (subagents). The same prompt and chunks, judged by Claude Code subagents. Several independent passes run under distinct labels (claude-a, claude-b, ...).
Runs as: handoff. Status: ready, conformance passed. - Codex judge. judge_prompt.md over each chunk, on a different model family from the Claude judges, so a certification is two families agreeing.
Runs as: self-driving. Status: needs a local requirement. - Reference judge (no model). Word-overlap check of each citation step against the abstracts it cites, plus a check that the experiment names a comparison. Exercises the judge contract with no model; a floor, not an auditor.
Runs as: self-driving. Status: ready, conformance passed. - theoria (arithmetic and derivation verifier). The upstream solve-and-audit pipeline over the computational claims extracted from candidates. Certifies or declines each claim, not each hypothesis.
Runs as: external. Status: designed, not yet built (takes claims, not candidates; runs through judge/theoria_adapter.py and is not yet under this contract).
Add an engine
An engine is a directory with a manifest and an adapter. The manifest records where the engine comes from and at which commit, its licence, how it reaches the corpus, what it needs to run and what it costs. The adapter has one entry point that turns a question into an answer; the hub builds the prompt, serves the corpus, logs every corpus call and checks every citation.
engines/my-engine/engine.yaml
engines/my-engine/adapter.py def generate(req) -> str | dictAn engine counts as built once it passes conformance. Any agent that speaks the Model Context Protocol can read a pack's corpus without an adapter of its own.