Your inputs seed a shared claim graph. Four stages (over 8 agents) read from it and commit back to it — initializing graph structure, verifying evidence, enriching concepts, and expanding new hypotheses.
You provide a claim, your topic priorities, and your expertise and interests. Those seed a shared claim graph; every stage reads from it and commits its results back.
Your expertise also steers the two generative stages, so what's mined and proposed stays close to your interests.
The Extractor reads your claim and seeds the graph — turning prose into structured nodes and causal edges, then committing them as the graph's first components.
Three independent judges weigh each edge against the retrieved evidence — a Coherence Judge (entailment vs. contradiction), a Methods Appraiser (design and bias risk), and a Causal Evidence Verifier (open risks and qualifiers). Their separate findings combine into one deterministic verdict, committed back to the graph unless the evidence falls short of what the verdict would claim.
A Concept Miner surfaces outcome concepts the graph implies but doesn't yet name; a Canonicalization Adjudicator merges duplicates so the vocabulary stays clean. The mined concepts are committed back to the graph.
A Hypothesis Proposer suggests new causal links. Independent Novelty Judges score each for novelty and saturation, combined by median; an eligibility gate (Gate_h) and a final ranking then decide which candidates survive, and the top survivors are committed — on your confirmation — as new hypotheses.
Every proposed hypothesis runs a three-stage gauntlet before any reach you. The thresholds and weights below are the configured defaults — echoed in each run's audit, not fixed truths.
Each scored 0–1. Most blend two ingredients — the proposing AI's opinion and hard semantic similarity measurements.
Plus four judgments the proposing AI gives directly, no quantified measurements blended in: plausibility, value of testing, importance, mechanism specificity.
Eligible only if all four hold at once. Missing any single bar disqualifies the hypothesis no matter how high the others are.
A panel separate from the proposer grades two things relative to the field, not just this graph. Numbers combine by the median — resisting a single outlier reviewer.
Establishedness is a 0–1 rating of how strongly the literature already treats the relationship as a known result.
This helps to re-order to demote the familiar (non-novel) ideas. Eligible hypotheses sort by this score, the higher the better.