Skip to content
Longevity Record

How every grade is produced

Methodology

This page describes exactly how a grade on this site is produced. It is written so that a reader can disagree with us specifically rather than generally — if you think a grade is wrong, this tells you which input to argue about.

The central rule

We grade claims, not interventions. A compound can have decent evidence that it shifts a blood marker and no evidence at all that it helps anyone live longer or better. Those are separate claims, they get separate grades, and we never average them into a single score.

This is not a stylistic preference. Presenting “NAD+ levels increased” as evidence for “lives longer” is the single most common error in longevity marketing. Our software checks for it: where a claim is about lifespan or healthspan and no linked human study measured that outcome, an automatic notice appears on the page.

What we record about each study

Bibliographic details — title, journal, year, authors, DOI — are retrieved from PubMed and never typed by hand, so a citation on this site cannot drift from its source. On top of that, an editor records the appraisal facts:

  • Subject. Human, animal, cell culture or modelling.
  • Design. From meta-analysis of randomised trials down to case series and mechanistic work.
  • Sample size and duration. How many participants, followed for how long.
  • Outcome type. Whether a hard outcome or a surrogate marker was measured.
  • Direction. Whether the result supports, contradicts or finds no effect for this specific claim.
  • Risk of bias. Assessed against standard domains; recorded as low, some concerns or high.
  • Funding and conflicts. Whether an interested party funded or authored the work.
  • Retraction status. Retracted work is excluded from grading and shown as excluded.

How a grade is derived

Those recorded facts are run through a fixed rubric that produces a suggested grade. The same inputs always produce the same suggestion — it is code, not impression, and the reasoning is printed on every claim under “why this grade”. The rules that do most of the work:

  1. 1.No human study, no human grade. If nothing has been tested in people, the best available grade is preclinical only, whatever the animal results show.
  2. 2.Surrogate markers cannot reach “strong”. If every human study measured a laboratory marker rather than a clinical or functional outcome, the grade is capped below strong however large the trials were.
  3. 3.Replication is required for the top grade. A single impressive trial is early evidence, not strong evidence.
  4. 4.Credible disagreement produces “mixed”. Where good trials point both ways, we say so rather than picking the flattering one.
  5. 5.Absence of evidence is labelled as such. “Insufficient evidence” means nobody has properly looked. It is not a judgement that something does not work — that is a separate grade, evidence against, and it requires good trials finding no effect.

An editor may override the rubric, but the override, the original suggestion and a written rationale are all published on the claim. You can always see where human judgement departed from the rules.

The grades

Strong human evidence
Consistent findings from multiple well-conducted human trials, or meta-analysis of randomised controlled trials, with adequate sample sizes and a clinically meaningful endpoint.
Moderate human evidence
More than one human trial pointing the same way, but limited by sample size, duration, risk of bias, or reliance on surrogate endpoints.
Early human evidence
One or a small number of small, short or preliminary human studies. Directionally interesting, not yet dependable.
Mixed evidence
Human studies disagree, with credible trials on both sides, or results reverse under better methodology.
Preclinical only
Evidence comes from animals, cell cultures or modelling. No human trial has tested this claim. Animal lifespan results do not establish human benefit.
Insufficient evidence
Too little credible research exists to judge the claim either way.
Evidence against the claim
Well-conducted human research indicates the claimed effect does not occur, or is too small to matter.

Outcome types

What a study measured determines what it can support. We record this explicitly so a biomarker result is never quietly upgraded into a health claim.

Lifespan
Death from any cause was measured.
Healthspan
Years lived free of major disease or disability were measured.
Disease outcome
A diagnosed condition or clinical event was measured.
Physical function
A measured capability such as walking speed, grip strength or VO₂ max.
Surrogate biomarkerSurrogate
A laboratory marker measured as a stand-in for health. A change here does not by itself demonstrate a health benefit.
Safety
Adverse events, tolerability or harm were measured.

Study designs and their weight

Study designs and the relative weight each carries in the rubric
DesignWeightHuman
Meta-analysis of RCTs10Yes
Systematic review8Yes
Randomised controlled trial7Yes
Non-randomised trial5Yes
Prospective cohort4Yes
Retrospective cohort3Yes
Case-control3Yes
Cross-sectional2Yes
Case series1Yes
Mechanistic study1No
Animal study1No
Cell / in vitro study0No
Modelling study0No

What we will not do

  • Publish a numeric score such as “83/100 effective”. The underlying evidence does not support that precision, and it invites false confidence.
  • Let commercial availability influence a grade. Whether a product is sold by our owner has no input into the rubric.
  • Publish an imported or machine-drafted summary without a named human reviewer approving it.
  • Describe a page as medically reviewed unless a named reviewer has genuinely reviewed it.
  • Adjust publication dates to appear fresher than we are.

Found something wrong? Our corrections policy explains how we handle it, and every grade change is recorded publicly.