ZACH Insights · Understanding evidence

How strong is
the evidence?

An illustrated guide to the evidence hierarchy—and the questions that turn a study result into a careful conclusion.

By Zacharias Razvi · Updated 3 October 2026
General education · Read from top to bottom, or follow the four sections below.

01 · See the structure

A hierarchy helps.
The question comes first.

“Research shows…” can refer to very different things: a laboratory experiment, an observation of thousands of people, a randomised trial or a review of many trials. They do not provide the same kind of answer.

The hierarchy shown here is a simplified guide for questions about the benefits of an intervention: does a treatment or programme improve an outcome compared with an alternative? Moving upwards usually means stronger opportunities to compare fairly, challenge alternative explanations and examine whether results repeat.

The position is a starting point, not a verdict. A well-conducted cohort study can be more informative than a badly executed trial. A systematic review of weak or irrelevant studies does not become strong evidence because it sits at the top.

For a different question, the most useful design changes. Long-term cohorts can inform prognosis. Case-control studies can investigate rare outcomes. Qualitative research explores experiences and barriers. Laboratory work can explain mechanisms. Those questions do not all belong on one treatment-effect ladder.

A teaching hierarchy · Intervention benefit
  1. 01 / SYNTHESISE Systematic reviews
    & meta-analyses
  2. 02 / TEST A COMPARISON Randomised controlled trials
  3. 03 / FOLLOW OR COMPARE Cohort & case-control studies
  4. 04 / TAKE A SNAPSHOT Cross-sectional studies
  5. 05 / IDENTIFY A SIGNAL Case reports & case series
Mechanisms & expert judgement Biological plausibility and informed interpretation matter. They do not, on their own, establish a treatment's benefit in people.
Original teaching illustration, not a formal grading system. Width does not represent study size, prevalence or a numerical quality score. Read the levels together with their explanations below.

Keep three questions separate. What type of study is it? How well was it carried out? How relevant is it to this decision? The hierarchy mainly helps with the first.

Follow the study design

Two groups. One fairer comparison.

A simplified parallel-group randomised trial. Select a stage to see what the design helps establish and what can still go wrong.

N R A B Δ Δ Participants Randomise Intervention Comparison Measure both

Participants

Eligibility defines who enters the trial. A random sample of a population and random allocation within a trial answer different questions. Representativeness is not guaranteed by randomisation.

Original design schematic. N, A, B and Δ are labels, not measured results. Moving lines indicate the reading order, not people changing groups. Blinding and several practical steps are omitted; read the bias and certainty discussion below.

02 · Understand the levels

Different designs.
Different strengths.

The purpose of a design is to make a particular comparison informative. Knowing its name helps you see what the researchers could control—and what could still explain the result.

01

Systematic review
& meta-analysis

A systematic review asks a defined question, searches for relevant studies using explicit methods and appraises what it finds. A meta-analysis is the statistical combination of compatible results. A review may include one, several or no meta-analyses.

Bringing studies together can improve precision and show whether findings agree. Combining numbers cannot remove bias in the original studies. Important differences in participants, interventions or outcomes may make one pooled estimate misleading. [1]

Look for: a transparent search, appropriate inclusion criteria, appraisal of bias, differences between studies and a conclusion that matches the evidence.

02

Randomised
controlled trial

A random process assigns participants to an intervention or a comparison group. This limits systematic differences caused by the way treatment is chosen. It helps address confounding, including factors the researchers did not measure.

Chance does not guarantee identical groups. Dropout, deviations from the assigned intervention, outcome measurement and selective reporting can still distort results. Random allocation is also different from randomly sampling a whole population: a trial can have a fair comparison while its participants remain quite specific. [2]

Look for: how allocation was performed and concealed, what the comparison received, missing results and the outcomes originally planned.

Illustration · Why a comparison matters
Eligible participants
A defined group, with a defined question
Random allocation
Chance determines the assigned group
Intervention Measure the chosen outcome
over a specified period.
Comparison Measure the same outcome
over the same period.
Compare outcomes between the groups.
An original schematic of a simple parallel-group trial. Improvement within one group is not enough: symptoms may change over time, expectations may help and other care may contribute. The between-group comparison helps distinguish those explanations. Real trials require further checks for bias.
03

Cohort &
case-control

A cohort study follows a defined group over time, comparing outcomes according to exposures such as physical activity. A case-control study starts with people who have an outcome and a suitable comparison group, then examines prior exposures. It can be useful for rare outcomes.

Researchers observe rather than randomly assign the exposure. Groups may differ in health, income, smoking or other factors. Statistical adjustment can address measured differences, but may leave residual confounding, measurement error or selection problems. An association needs careful causal interpretation. [3]

Look for: how groups were selected, the timing of exposure and outcome, measurement quality and credible alternative explanations.

04

Cross-sectional
study

A cross-sectional study measures a group around one time point. It can describe how common a characteristic is and which measurements vary together. It is useful for a snapshot of a population.

Timing is the central limitation. If activity and pain are measured together, did less activity contribute to pain, did pain reduce activity, or did something else influence both? A snapshot cannot settle that sequence by itself.

Look for: who was included, whether the sample represents the intended population and whether the headline incorrectly turns a same-time association into a causal claim.

05

Case report
& case series

A case report describes one person or event; a case series describes several. Such reports can highlight something unexpected, suggest a new research question or draw attention to a possible harm.

Without an appropriate comparison, they cannot reliably show what would have happened without the intervention. A striking personal improvement does not establish its cause or tell us how often others will improve.

Look for: a useful observation, while keeping the conclusion proportionate. A signal is a reason to investigate, not proof of a general treatment effect.

+

Mechanisms
& expertise

A plausible biological pathway can explain why an intervention might work. Cells, tissues and controlled experiments help build that explanation. Clinical expertise helps interpret evidence in context.

However, changing a molecule is not the same as improving everyday function. A short-lived biomarker response does not establish a long-term health benefit. Mechanisms and expertise contribute to understanding; direct outcome research tests whether the expected benefit occurs.

Look for: the link between the proposed mechanism, the actual measurement and the outcome people care about.

03 · Beyond the design label

How much confidence
should the result earn?

A study can be well designed in principle and weak in execution. A body of evidence can contain several studies yet still leave an important question unresolved. This is why evidence appraisal goes beyond counting papers or ranking their titles.

GRADE is an established framework for assessing certainty in a body of evidence for a specific outcome. It uses four categories: high, moderate, low and very low. Certainty can differ between benefits and harms within the same review. The following are five important reasons confidence may be reduced. [4]

  • Risk of bias Could the way a study was conducted or reported systematically distort its result?
  • Inconsistency Do studies disagree substantially, and can the differences be explained?
  • Indirectness Are the people, intervention, comparison or outcome different from the question being answered?
  • Imprecision Is the estimate uncertain enough that materially different conclusions remain plausible?
  • Publication bias Could missing or selectively published results make the available evidence look more convincing than it is?

This guide explains appraisal concepts; it does not perform a formal GRADE assessment. A recommendation also considers benefits, harms, burden, feasibility and what matters to the person making the decision. Certainty and recommendation strength are related, but distinct. [5]

04 · Make the result understandable

Size. Uncertainty.
Relevance.

How large is the difference?

A relative percentage sounds impressive without a starting point. A reduction from 10% to 8% is a 20% relative reduction, but a difference of 2 percentage points: two fewer events per 100 people over the stated period.

If the starting risk were 1%, the same 20% relative reduction would take it to 0.8%—a difference of 0.2 percentage points. Always ask for the absolute numbers, the comparison and the timeframe. [6]

For an outcome such as pain, ask about the measurement scale and whether the difference is meaningful in daily life. “Statistically significant” does not tell you that the benefit is large, important or certain to occur for you. [7]

Illustration · Invented one-year example
Comparison group 10 in 100
Intervention group 8 in 100
0% 50% 100%
20% Relative reduction
(10 − 8) ÷ 10
2 points Absolute difference
10% − 8%
Both bars use a 0–100% scale. These invented numbers illustrate arithmetic, not a real treatment or your personal risk. They assume comparable groups and the same one-year period.

How uncertain is the estimate?

An estimate belongs with an uncertainty measure, commonly a confidence interval. A narrow interval indicates greater statistical precision; a wide interval leaves a broader range of effects compatible with the data and model. Precision does not remove systematic bias. [6]

Read the interval against meaningful benefit and harm—not only whether it crosses a “no difference” value. A result described as non-significant may still be too imprecise to rule out an important effect. Conversely, a very precise difference can be too small to matter.

Does the finding answer this question?

Keep the conclusion close to the people, dose, comparison, outcome and follow-up actually studied. Treating existing symptoms is different from preventing new cases. Improved test performance is different from fewer diagnoses years later. An average group effect is not a promise for every participant.

  1. What was the precise question?

    Name the population, intervention or exposure, comparison, outcome and time period.

  2. Was the design suitable—and well carried out?

    Use the hierarchy to orient yourself, then examine the methods and possible sources of bias.

  3. What changed, by how much, and with what uncertainty?

    Read the absolute result, its interval, adverse outcomes and the practical meaning of the scale.

  4. How does it fit the wider evidence and this decision?

    Look for replication, a trustworthy synthesis and a realistic match to the intended person and setting.

A careful conclusion names its limits. Evidence does not become less useful when uncertainty is made visible. It becomes easier to use responsibly.

Sources & further reading

Read the methods
behind the guide.

The illustrations and examples are original educational material. The hierarchy is a simplified model for intervention-benefit questions, rather than a reproduction of a formal evidence classification.

  1. Cochrane Handbook, chapter 10: Analysing data and undertaking meta-analyses. Synthesis, heterogeneity and why pooling can mislead.
  2. RoB 2 authors: Risk of bias in randomised trials. The official tool and guidance for trial appraisal.
  3. CDC Field Epidemiology Manual: Designing and Conducting Analytic Studies in the Field. Cohort, case-control and observational study principles.
  4. Cochrane Handbook, chapter 14: Summary of findings and certainty of evidence. Outcome-specific certainty and the GRADE domains.
  5. GRADE Working Group: Official guidance and publications. Certainty assessment and the distinction between evidence and recommendations.
  6. Cochrane Handbook, chapter 6: Choosing effect measures. Absolute and relative effects and measures of uncertainty.
  7. Cochrane Handbook, chapter 15: Interpreting results and drawing conclusions. Practical meaning, uncertainty and application.

Sources checked 3 October 2026. This page provides general research literacy, not an appraisal of a particular treatment or individual medical advice.

ZACH / Connect the mechanism to the decision

What can you use this for?

01 Question
02 Study design
03 Estimate & uncertainty
04 Relevance

A reading sequence, not a measured causal model.

Keep the mechanism precise.

A design creates opportunities for a useful comparison. Execution determines how well those opportunities were used. Certainty also depends on consistency, precision and whether the participants and outcomes answer the question you want to ask.

A concrete way to apply the idea.

Write the claim in one sentence. Then record who entered the study, what was compared, the outcome, the follow-up and the uncertainty. Make a second sentence that says only what the study supports. The difference between the two sentences reveals where interpretation has exceeded the evidence.

What we do not know.

A hierarchy cannot substitute for a methods appraisal. Statistical significance does not establish clinical importance, and a pooled average does not forecast every individual response. A review inherits limitations from its included studies.

Cochrane · Handbook for Systematic Reviews of Interventions What this source supports: Methods for evaluating intervention evidence, bias and the meaning of pooled findings.

Continue learning / Follow the question

For a structured self-guided format, explore ZACH Learning. Check its current availability and scope before choosing a programme.