How we rate evidence

WhatEvidence.Health
Evidence Rating System

Evidence Rating System

This page explains how WhatEvidence.Health rates the strength, reliability, and practical meaning of health evidence. The goal is simple: separate reliable human evidence from early signals, uncertainty, and unsupported conclusions.


Bottom Line

WEH ratings are based on study type, sample size, human relevance, direct measurement, risk of bias, consistency, and real-world usefulness. Stronger ratings require reliable human evidence, preferably systematic reviews, meta-analyses, and randomized controlled trials with adequate sample sizes. As a general rule, WEH treats N ≥ 50 as the minimum preferred sample size for a trial to meaningfully influence the rating.

Why This System Exists

Health evidence is not just “good” or “bad”. Some evidence is strong and directly useful. Some is promising but early. Some is biologically plausible but not proven in humans. Some is too weak, too small, or too indirect to guide decisions.

WEH uses clear rating terms so readers can quickly understand how much confidence the evidence deserves.

WEH Rating Terms
Supported
Reliable evidence supports the finding

This rating is used when consistent human evidence supports the finding. It usually requires higher-quality evidence such as systematic reviews, meta-analyses, or randomized controlled trials. The result should be directly measured, reasonably consistent, and not explained mainly by bias or chance.

Promising
Encouraging, but not definitive

This rating is used when early or moderate evidence is encouraging, but more rigorous trials are needed. The biological mechanism may be plausible and some human data may exist, but the evidence may be limited by small studies, indirect outcomes, inconsistent methods, or incomplete long-term data.

Uncertain
The evidence is mixed or insufficient

This rating is used when the evidence does not allow a confident conclusion. Studies may be too small, too different from each other, poorly controlled, or focused on indirect outcomes. “Uncertain” does not mean disproven. It means the evidence is not strong enough yet.

Not Supported
Reliable evidence does not support the finding

This rating is used when better-quality evidence does not show a meaningful effect, or when the available evidence is too weak to support the conclusion being presented. It may also apply when the proposed mechanism sounds plausible but human outcomes do not confirm it.

GRADE

GRADE is a plain-language way of showing how confident we are in the evidence behind a finding. WEH commonly uses terms such as High, Moderate, Low, or Very Low.

A higher GRADE means the result is more likely to be reliable. A lower GRADE means the finding may change when better studies become available.

Study Types We Weigh More Heavily
  • Systematic reviews: structured reviews that collect and assess multiple studies on the same topic.
  • Meta-analyses: reviews that statistically combine results from multiple studies.
  • Randomized controlled trials (RCTs): studies where people are randomly assigned to an intervention or comparison group.
  • Controlled human studies: studies in people with a comparison group and clearly measured outcomes.
The N Rule

N means the number of participants in a study.

For WEH ratings, a trial with N ≥ 50 is generally the minimum preferred size for meaningfully influencing the rating. Smaller studies can still be discussed, but they are treated cautiously because small studies are more likely to exaggerate effects, miss safety signals, or produce unstable results.

A large study is not automatically good, and a small study is not automatically useless. But sample size matters.

How WEH Interprets Evidence
1. Human Evidence Comes First

WEH gives much more weight to studies in humans than to laboratory, cell, or animal studies.

Lab and animal research can explain mechanisms and generate ideas, but it does not prove that the same effect occurs in real people at real-world doses.

2. Direct Outcomes Matter

WEH looks for outcomes that actually answer the question being studied.

For example, a study measuring a meaningful clinical outcome is usually more useful than a study measuring only a small change in a marker that may or may not matter in daily life.

3. Mechanism Is Helpful, But Not Enough

A mechanism explains how something might work biologically.

Mechanisms are useful, but WEH does not treat mechanism alone as proof. A pathway can make sense in theory and still fail to produce a meaningful human result.

4. Caveats Are Part of the Verdict

Caveats are the important limitations that stop a result from being overinterpreted.

  • Was the study too small?
  • Was the duration too short?
  • Was the outcome indirect?
  • Were the results inconsistent?
  • Was the product, dose, or population different from real-world use?
5. Risk of Bias Matters

Risk of bias means the chance that a study’s result is distorted by its design, funding, methods, missing data, selective reporting, or lack of blinding.

WEH commonly describes risk of bias as Low, Low–Moderate, Moderate, or High.

6. Consistency Matters

One positive study is rarely enough. WEH looks for consistency across studies.

A finding becomes more convincing when different research groups, different populations, and different study designs point in the same direction.

7. Safety Is Considered Separately

A result can look promising and still have safety concerns.

WEH separately considers side effects, medication interactions, pregnancy and breastfeeding concerns, dose uncertainty, product quality, injectable use, and long-term safety gaps.

Evidence Summary Table Format

On mobile: swipe sideways to view all columns.

Table Term Meaning How WEH Uses It
Citation The study or review being assessed. Allows readers to trace the evidence back to the original source.
Study Type The design of the research, such as RCT, systematic review, meta-analysis, cohort study, or lab study. Higher-quality designs generally receive more weight.
N The number of participants. N ≥ 50 is the usual minimum preferred size for a trial to meaningfully influence the rating.
Key Result The main finding reported by the study. WEH checks whether the result is meaningful, direct, and relevant.
Verdict The WEH rating: Supported, Promising, Uncertain, or Not Supported. Summarises the practical strength of the evidence.
Risk of Bias The likelihood that the result is distorted by study design or reporting problems. Higher bias lowers confidence in the result.
What Can Lower a Rating
  • Small sample size, especially N < 50
  • No control group
  • No randomization
  • No blinding
  • Short follow-up
  • Indirect or weak outcomes
  • Large dropout rate
  • Selective reporting
  • Industry funding without strong safeguards
  • Results that are not repeated by other studies
What Can Raise a Rating
  • Multiple well-designed human trials
  • Systematic reviews or meta-analyses of good-quality studies
  • Consistent results across studies
  • Directly measured outcomes
  • Adequate sample sizes, preferably N ≥ 50 per meaningful trial
  • Low risk of bias
  • Clear dose, timing, population, and comparison group
  • Meaningful real-world benefit
  • Transparent safety data
Final Rule

WEH does not rate evidence by confidence, popularity, or biological plausibility alone. A higher rating requires reliable human evidence, adequate study size, direct outcomes, low or acceptable risk of bias, and a result that matters in real life.

WhatEvidence.Health · Evidence Rating System

This is general health information, not personal medical advice. Consult your doctor for personal medical decisions.