Piper by Focena

Rare disease has too little evidence — and too much confidence.

Piper reads the literature the way a careful reviewer would. Three frontier models extract each paper independently and then argue about it. Disagreement is kept as signal rather than averaged away. Every claim carries a ceiling on how much certainty it is allowed to assert — and nothing is ever stated above the certainty it has earned.

What has been built

A disciplined substrate, not a demo

These are counts of work already done — papers appraised, panels run, measurements extracted and made comparable across genes.

3,275
papers appraised

Each read by a three-model adversarial panel, not summarised by one.

144,625
measurements extracted

Physical values with their units, sample size, assay and direction — not sentences about them.

3,112
disease genes spanned

The substrate is gene-general by design; a tool that only works for one disease has failed.

2,969
adversarial panel runs

Three frontier models extract independently, then challenge each other across rounds.

2,629
papers yielding measurements

A paper enters the lake only where it reports a number we can identify and compare.

50
evidence corpora

Assembled around anchor diseases, mechanisms and delivery routes.

Derived from the research substrate on 28 August 2026 · framework 0.43.2. These are counts of our own work — papers read, panels run, measurements extracted — never a republication of any source’s content.

Our mission

To make rare-disease evidence legible and honest — so that the people deciding where scarce research money goes can tell what is known from what is merely said.

A disease with a few hundred known patients does not get a large, well-powered literature. It gets a handful of small studies of wildly varying quality, a great deal of hope, and tools that will confidently summarise all of it into a paragraph that sounds settled. For a foundation with one grant to give or a researcher with one experiment to run, that confidence is the problem, not the help.

Piper is built the other way round: to be careful about what it does not know, to keep the reasoning that produced every judgment, and to refuse rather than reach when the evidence will not carry the weight.

How it works

Four steps, and the argument is the product

The extracted value is almost a by-product of the reasoning that produced it. That reasoning is kept, attributed, and inspectable.

Appraise

Three frontier models — from three different vendors — read each paper independently and answer the same set of methodological questions a careful scientist would ask. Was it big enough? Could bias have crept in? Was it independently repeated?

Argue

The models then evaluate each other and are pushed on what they still disagree about. Easy agreement is treated as an anti-signal, not as truth: three models trained on overlapping data can be confidently wrong the same way.

Measure

Separately, the paper’s own numbers are extracted — value, units, sample size, assay, the system it was measured in, and the direction it moved — and given an identity precise enough that the same measurement in two papers can actually be compared.

Arbitrate

Where the panel splits, a ticket goes to a human reviewer. A human ruling is the only path by which anything becomes confirmed. Machine output never promotes itself.

The discipline

Most tools optimise for coverage. This one optimises for trustworthiness.

Five commitments, each structural rather than aspirational — they are enforced in the code and the database, not promised in a policy.

Commitment

Machine consensus never impersonates human judgment

Everything the panel produces is presumptive and stays that way until a person rules on it. There is no threshold at which agreement becomes truth, and no automatic promotion.

Commitment

Disagreement is preserved, never averaged

Reviewer and model disagreement is recorded as structure. Averaging it away would destroy the one signal that tells you where the evidence is genuinely contested.

Commitment

Nothing goes missing quietly

At every step the arithmetic must close: everything that entered either came out or is accounted for by a named, true reason. A step that cannot account for itself fails loudly rather than silently dropping rows.

Commitment

Identifiers are looked up, never recalled

Gene, ontology and chemical identifiers are resolved against pinned reference indexes rather than generated from a model’s memory. A confabulated identifier that happens to exist is indistinguishable from a correct one downstream — so the label is the witness and the code is derived from it.

Commitment

Regulatory approval is not ground truth

A standing test compares two well-known trials of the same disease: the larger, more rigorous one that failed cleanly, and the smaller, thinner one that was approved. The framework must score the rigorous trial higher on methodological quality. If approval ever starts driving credibility, that test fails and the build stops.

Scope

Gene-general by design

The anchor disease is CASK-related disorder — a few hundred known patients worldwide — but the substrate spans thousands of disease genes on purpose. The point of a general evidence base is to carry what is known about well-studied genes onto the ones nobody has studied.

Who it is for

Built for the people who have to choose

Access is for professional and research use. Foundations use it to decide what to fund; researchers to decide what to run next.

Rare-disease foundations

Where the next research dollars would do the most good — with credibility and leverage kept visibly separate, so a trustworthy finding and a high-impact one are never confused. Includes plain-language material a foundation can put in front of its own community.

Researchers

What is already measured, what is genuinely open, and which single experiment would settle the most at once. Experiments are ranked by how cheaply they can kill a hypothesis, never by how likely it is to be true.

Clinicians and geneticists

The evidence behind a gene or variant assembled in one place, with the open questions named. The variant panel lays out the evidence and shows how the standard clinical rules would weigh it — and deliberately classifies nothing. That call is yours.

Behind the sign-in

The deep research

Members reach the working surfaces over the whole substrate. Every one of them carries its own certainty ceiling and shows its sources.

Research-leverage brief

Where the next research dollars could go, credibility and leverage kept apart on purpose.

Findings

Per-paper credibility summaries — including how each of the three models read it and exactly where they split.

Variant panel

Assembles the evidence for a gene and variant and shows how the standard clinical rules combine it. Classifies nothing.

Mechanism and repurposing

Drugs studied for diseases that share a mechanism with yours — hypotheses to investigate, never treatments.

Research agenda

The honest unknowns the framework flagged, and the fundable study directions that would close them.

Therapeutic strategies

The main approach classes scored on two gates at once: can it work, and can it reach the brain.

The limits

What this is not

A tool whose whole claim is epistemic honesty has to be honest about itself first.

  • It is not medical advice, and it diagnoses nothing. It does not make clinical, genetic or variant-pathogenicity determinations, and nothing it produces should be acted on alone.
  • Today, essentially all of its content is provisional. Human arbitration trails the framework by design — it never blocks the analysis — which means the queue of machine-generated findings runs well ahead of what a person has ruled on. Provisional items are labelled as such everywhere they appear, and you should read them as working hypotheses.
  • Its depth is uneven. Coverage is deepest around the anchor disease and the neurodevelopmental genes nearest to it, and thins out from there. The framework reports what it does not hold rather than filling the gap with something plausible.
  • It has not proven forecasting skill. The generative layers propose hypotheses; they are not validated predictors. Every hypothesis therefore ships with the cheapest experiment that would falsify it, which is what makes it safe to act on at all.
  • Nothing reaches a patient through us. Everything here is upstream of the bench. Work is vetted by actual laboratory and clinical science before it means anything for a person.

Access is by invitation

Piper is in a controlled launch with a small number of partner organisations. If your foundation, lab or clinic works on a rare genetic disease, tell us what you are trying to decide.

Research aid only — not medical advice. Piper is a research and educational aid. It is not medical advice, not a diagnosis, and not a clinical determination. Its output is hypothesis-generating and may be incomplete, provisional, or wrong, including AI-generated errors. Always consult your own qualified healthcare professional before making any medical or treatment decision.

Piper by Focena — calibrated evidence for rare disease