Understanding evidence levels: what each study design shows and what it does not
Evidence levels grade confidence, not truth: what randomised, observational and animal studies support, how GRADE rates them and where the pyramid breaks.
Scientific reviewPosition not yet filled, stated openly.
10 min readLast substantive review
Open access
Executive summary
An evidence level does not judge whether a statement is true. It judges how much confidence the available data deserve. The rating is given per outcome, not per study: in the same review, the cure rate can count as moderately certain while bone density counts as very uncertain. What matters is therefore not the label on the study design, but which alternative explanation that design can actually rule out.
19primary sources
58 %of them level 1 to 2
4species studied
2007–2025publication years
Key points
GRADE has four levels of certainty and assigns them per outcome: in a 2023 review the cure rate counted as moderately certain, while bone density in the same review counted as very uncertain.
Across an analysis of 1,583 meta-analyses, randomised and observational studies arrived on average at practically the same result: the design label alone says little about reliability.
Missing blinding enlarges effects measurably: by roughly a tenth in clinical trials, and by 59 per cent on average in a pooled analysis of ten animal experiments with mostly subjectively judged outcomes, where three outliers carry that average.
A strong recommendation is not strong evidence: of 131 strong recommendations drawn from 133 guidelines surveyed, 104 rested on low or very low certainty.
In equine medicine the data base is thin: a 2024 systematic review screened 7,009 publications and kept nine of them.
What an evidence level measures: confidence, not truth
An evidence level is a statement about the state of knowledge, not about nature. GRADE, the system adopted by numerous professional bodies worldwide, distinguishes four levels: high, moderate, low and very low certainty. These levels answer a single question: how likely is it that further good studies will still shift today’s result? A consensus article by an international methods group set out this framework in its current form in the British Medical Journal in 2008. It separates two things that were previously mixed up on a regular basis: the certainty of the evidence and the strength of any recommendation derived from it.
The most important point, and the one most often missed: certainty is assigned per outcome, not per study and not per article. A 2023 Cochrane review of parathyroid surgery pooled eight randomised trials with 447 adults. For the outcome of cure, the authors rated certainty as moderate; for bone density at the lumbar spine, in the same piece of work, as very low. Anyone who reads only the title and the conclusion will take both to be equally well supported.
Why do two specialists grade the same study differently?
Because every downgrade calls for a judgement. How large does an imprecision have to be before it counts? Is a contradiction between two studies real, or is it down to chance? GRADE makes these judgements visible and requires them to be justified, but it does not make them for anyone. Two careful teams can therefore reach different levels from the same body of data.
The levels, and what each of them rules out
The classic ranking is usually drawn as a pyramid: case reports and laboratory experiments at the bottom, observational studies above them, then randomised trials, and systematic reviews at the top. This order is not a ranking of importance. It sorts designs by which alternative explanation each one can rule out. A case report rules out nothing; it generates a hypothesis. Allocation by lot is the only procedure that also rules out unknown confounders, because it pushes chance in between the groups.
What each design can rule out, what stays open, and the short circuit that typically follows.
Design
Rules out
Stays open
Typical false conclusion
Case report, case series
nothing; the design describes a course of events
whether the course would have been the same without treatment
“It worked”
Laboratory experiment, cell model
interference from the rest of the organism
whether the mechanism matters in a living animal or human
reading a mechanism as proof of effect
Cross-sectional survey
nothing about the order in time
which came first, cause or consequence
reading an association as a cause
Cohort and case-control study
gross reversal in time, and some known confounders
unknown confounders, self-selection into the groups
“People who do X stay healthier, so X helps”
Randomised trial
known and unknown confounders at the start of the trial
whether the result holds for other animals, people and conditions
confusing efficacy under trial conditions with everyday benefit
Systematic review, meta-analysis
arbitrariness in the selection of studies
every weakness of the included studies
“It is a meta-analysis, so it is proven”
The level is therefore not a verdict on the quality of the researchers. A carefully controlled cohort study can be more reliable than a badly conducted randomised trial, and a meta-analysis built from ten weak papers stays weak. The pyramid ranks designs, not individual publications.
Randomisation: what it achieves and what it does not repair
How much bias a careless randomisation produces can be measured. An analysis of 2,443 randomised trials drawn from 228 meta-analyses in the Cochrane Library compared papers at low risk of bias with those at unclear or high risk. Trials without a traceable random allocation, without concealed group allocation or without blinding showed treatment effects that were on average some eight to thirteen per cent larger. The difference is not dramatic, but it is systematic and almost always in the same direction: in favour of the treatment under test.
Five things, on the other hand, randomisation does not repair: a group that is too small, an outcome nobody cares about, an observation period that is too short, a heavily selected population, and selective reporting of the results. These five points are precisely why GRADE can downgrade a randomised trial far enough for it to land on the same certainty level as an observational study.
In veterinary medicine, reporting on these points has only partly caught up. A review of the 2009 and 2019 volumes of a journal of veterinary anaesthesia found that full description of the randomisation rose from 13.6 to 85.7 per cent, and the reporting of a sample size calculation from zero to 20 per cent. Blinding stayed at around half of the papers and did not move in ten years. Not a single item reached the mark the authors had set as their target.
Observational studies are not automatically weaker
The pyramid suggests that observational data are fundamentally less reliable. A 2014 Cochrane methodology review examines exactly that: fourteen methodological reviews holding data from 1,583 meta-analyses across 228 conditions compared the effect estimates of randomised and observational studies addressing the same question. On average they did not differ meaningfully, and eleven of the fourteen reviews found no difference at all. The authors did not conclude from this that design is irrelevant, but that factors other than the label explain the discrepancies.
For some questions the observational study is in any case the only possible design: rare harms, long time spans, situations in which allocation by lot would not be defensible. A 2025 review therefore proposes replacing the pyramid with an amphitheatre for questions of post-authorisation medicine safety. All data sources stand in a circle around the question, and none automatically stands on top.
Does that mean I can read observational data like randomised data?
No. The finding holds for the average across many questions, not for the single study in front of you. With an observational study, what remains decisive is whether the groups being compared already differed before treatment, and whether that difference was balanced out. If nothing at all is reported on this, the result cannot be interpreted.
Whoever measures has a say: the problem with soft outcomes
The softer the outcome, the more the expectation of the person doing the measuring comes into play. A pooled analysis of ten animal experiments with 2,450 animals in total compared blinded and unblinded assessment of the same experiments. Unblinded assessment exaggerated the treatment effect by 59 per cent on average. In most of these experiments the outcomes were subjective, that is, observations rather than measured values. That average, however, is carried by three experiments with a very large deviation; without them, the remaining seven leave an exaggeration of about one fifth, which is no longer statistically significant.
In small animal medicine this mechanism has been measured directly. In the placebo group of a registration study of 58 dogs lame from osteoarthritis, owners saw an improvement in 39.7 per cent of cases, and the examining veterinary surgeons, scoring lameness, in 44.8 per cent. Over the same period the force plate showed no change in 46 of the 58 dogs. In cats with joint pain, five studies covering 96 animals found the same pattern: 54 to 74 per cent of the placebo animals counted as a success in the owners’ judgement, but only 10 to 63 per cent by objectively measured activity.
Why does a placebo work in an animal that knows nothing about it?
Because it is not the animal that supplies the outcome, but the human being who judges it. Add to that regression to the mean: treatment usually starts when the complaints are at their worst, and a better phase often follows anyway. The two together produce a visible improvement without any active substance at all.
GRADE in detail: five reasons to downgrade, three to upgrade
GRADE starts randomised trials at high certainty and observational studies at low certainty, then moves them. Five reasons lead downwards:
Risk of bias: no concealed allocation, no blinding, many dropouts without explanation.
Inconsistency: the studies point in different directions, with no way to explain it.
Indirectness: a different population, a different comparison, or a surrogate outcome instead of what actually matters.
Imprecision: too few events, and the result is compatible with benefit and with harm at the same time.
Suspected publication gaps: strikingly many small positive studies, and hardly any small negative ones.
Three reasons, conversely, allow observational data to be upgraded: a very large effect, a clean dose-response relationship, and confounders that would have shrunk rather than inflated the effect found. An observational study can thus rise to moderate certainty and a randomised one fall to very low. Within GRADE, then, the pyramid is not a fixed shelf but a starting position.
Certainty of the evidence and strength of a recommendation are two different axes, and they often come apart. A systematic survey of 133 guidelines carrying 433 recommendations on acupuncture found 303 recommendations with a GRADE rating. Of those, 131 were worded strongly, and 104 of these strong recommendations rested on low or very low certainty. Emphatic wording is therefore no evidence of emphatic data.
Where the pyramid breaks: reviews, species boundaries and thin evidence
In the classic drawing, the systematic review sits at the top. A much-cited proposal from 2016 takes it out of the pyramid and turns it into the lens through which the levels below are viewed. The reason is simple: a review generates no new data. It can only order and weight what is already there.
A stocktaking from the same year shows how necessary this correction is. Between 1991 and 2014 the number of systematic reviews published each year grew by 2,728 per cent, while all indexed papers together rose by 153 per cent. On antidepressants for depression alone, 185 meta-analyses appeared between 2007 and 2014. The author concluded that the great majority of the reviews being produced are redundant, misleading or shaped by vested interests.
In equine medicine the opposite problem is the rule: too little rather than too much. A 2024 systematic review of drugs against equine herpesvirus 1 screened 7,009 publications and kept nine; the risk of bias was moderate to high and the groups were small. A review of electrotherapies in horses, dogs and cats screened 5,385 publications, included 41 and arrived at an average of 3.7 papers per procedure, 20 of which carried a high risk of bias. The conclusions are correspondingly cautious: the electrotherapy team considers the evidence sufficient for not a single application in horses, dogs or cats, and the herpesvirus team saw at best a minimal or limited benefit for all the drugs examined.
The second fault line is the change of species. A systematic review compared six treatments for which a clear result existed in humans with the corresponding animal experiments. For corticosteroids after traumatic brain injury, the animal models showed a benefit and the clinical trials none. The drug tirilazad shrank the infarct area by 29 per cent in the animal model and improved neurological scores by 48 per cent, while stroke patients on the same drug fared worse than those without it.
Part of this discordance goes back to methods, not to biology. An overview of 31 systematic reviews of animal studies found that 29 per cent of the individual studies reported any randomisation at all, 15 per cent a concealed allocation and 35 per cent a blinded outcome assessment. A Swiss analysis of 1,277 animal experiment applications found the seven standard measures against bias described at all in only 2.4 to 19 per cent of the applications.
Placing a study report in five minutes
The following seven questions can be put to almost any report, usually to the abstract alone. They are no substitute for a formal appraisal, but they reliably separate the reports that show something from those that merely assert something.
Which species was studied: human, horse, another animal, cell culture? If it is not stated, that is already information.
Which design: allocation by lot, observation, laboratory experiment, review?
How many participants? Twenty animals and two thousand people do not carry the same weight.
Which outcome: a laboratory value, or something an animal or a person can actually feel?
Who did the assessing: a device, a blinded person, or someone who knew the treatment?
Over what period? Six weeks say nothing about two years.
Was there a comparison group? Without a comparison, any improvement is ambiguous.
What is left at the end of a systematic review
Original analysis
Figures taken directly from the methods and results sections of the five reviews cited on this page; the third column reports each author team’s own rating, not ours, and the fourth names the conclusion the work precisely does not support.
Review (field, year)
Starting pool and included studies
Certainty according to the authors
What does not follow from it
Drugs against equine herpesvirus 1 (horse, 2024)
7,009 publications screened, 9 included
risk of bias moderate to high, small groups
that the drugs examined are ineffective: too little has been tested to decide that
Electrotherapies in horses, dogs and cats (2022)
5,385 publications screened, 41 included, on average 3.7 per procedure
20 of the 41 papers at high risk of bias
that a single positive paper establishes a procedure
Nutritional supplementation in unexplained infertility (human, 2019)
5,663 records screened, 61 studies, 59 of them randomised
very low for every outcome
that many randomised trials automatically add up to high certainty
Animal-assisted therapy in dementia (human, 2019)
9 randomised trials, 305 participants, one study with horses
very low to moderate, depending on the outcome
that an effect on mood carries over to quality of life
Parathyroid surgery (human, 2023)
8 randomised trials, 447 adults
moderate for cure, very low for bone density
that one piece of work has a single certainty level valid for everything
Limitations and uncertainty
The average values cited come from methods research and apply to many questions taken together. They do not predict how strongly any one individual study is biased.
GRADE and the evidence pyramid are conventions of a professional community, not experimental findings. Applying them calls for judgements that vary between assessors.
The comparison of randomised and observational studies rests on methodological reviews of very uneven quality; only a minority of them were rated at low risk for selective reporting. Modern causal analysis methods were not represented in them at all.
The figures on the caregiver placebo effect come from dogs and cats. For the horse, this research found no comparable quantification.
Several key methodological papers date from 2007 to 2016 and have not been replaced by newer syntheses to date. Whether their figures still reflect current practice is an open question.
The veterinary methodological literature is thin: part of the evidence used here comes from anaesthesia and small animal medicine, not from equine medicine.
Open questions
How large is the caregiver placebo effect in the horse, where lameness is mostly scored by eye?
Do objective gait measurement systems actually reduce this effect in clinical trials, or do they merely shift it?
Does GRADE need adapting for veterinary questions when only three or four studies on a procedure exist worldwide?
Do automated evidence syntheses reduce the redundancy of reviews, or do they increase it?
Frequently asked questions
Is a randomised trial always better than an observational study?
No, but it answers a different question. Allocation by lot is the only procedure that also balances unknown differences between the groups, which makes it the strongest protection against false conclusions about causes. It does not protect against groups that are too small, short observation periods or ill-suited outcomes. In an analysis of 1,583 meta-analyses, randomised and observational results on the same question did not differ meaningfully on average. For rare harms and long time spans, the observational study remains the only workable design.
What does “very low certainty of evidence” actually mean?
It means that the true effect very probably lies clearly elsewhere than the reported result, in an unknown direction. The statement is therefore not a basis for a decision but a description of the state of research. One point matters here: very low certainty does not mean that the treatment does not work. It means that we do not know. In a 2019 review of nutritional supplementation for unexplained infertility, certainty sat at this level for every outcome, even though 59 randomised trials were included.
Why does a single study never sit at the top, even when it is well done?
Because a single study cannot check its own result. Chance, local peculiarities of the population and choices made during the analysis all play a part in every individual paper, and only replication by independent groups separates signal from noise. A well conducted single study can still be the best that exists on a question. It then remains the best available basis, but with low certainty, not with high.
Does a study in horses count as evidence for humans?
No, and no more so the other way round. A systematic review compared six treatments with a clear result in humans against the corresponding animal experiments and found several plain contradictions: one drug shrank the infarct area by 29 per cent in the animal model, while stroke patients on the same drug fared worse. An animal result justifies a hypothesis for the other species, nothing more. The testing has to take place in the target species and the target population.
How do I spot a weak study report in the media?
By what is missing. If the text names neither the species studied, nor the number of participants, nor a comparison group, nor the observation period, then there is no basis for placing it at all. A second warning sign is the confusion of association and cause: wording such as “lowers the risk” presupposes a design that can test causes. A third is the surrogate outcome: an improved laboratory value is not the same as improved wellbeing.
Why are methodological papers from 2007 or 2014 cited here?
Because nothing has replaced them on their particular question. The Cochrane methodology review comparing randomised and observational results still stands in its 2014 version; the compilation on agreement between animal experiment and clinical trial dates from 2007. In methods research a paper ages more slowly than in clinical research, because it examines procedures rather than substances. Where more recent figures exist, they have been given preference here.
Sources
Guyatt GH, Oxman AD, Vist GE, Kunz R, Falck-Ytter Y, Alonso-Coello P, Schünemann HJ. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ, 2008 (Other | Species-independent)DOI 10.1136/bmj.39489.470347.AD Introduces the GRADE framework and makes the case for a strict separation between the certainty of the evidence (four levels) and the strength of a recommendation.
Murad MH, Asi N, Alsawas M, Alahdab F. New evidence pyramid. Evidence-Based Medicine, 2016 (Other | Species-independent)DOI 10.1136/ebmed-2016-110401 Argues that systematic reviews and meta-analyses are not the top level of the pyramid but the tool with which the levels beneath them are read.
Anglemyer A, Horvath HT, Bero L. Healthcare outcomes assessed with observational study designs compared with those assessed in randomized trials. Cochrane Database of Systematic Reviews, 2014 (Systematic review | Human)DOI 10.1002/14651858.MR000034.pub2 Across 14 methodological reviews holding data from 1,583 meta-analyses on 228 conditions, the effect estimates of randomised and observational studies did not differ meaningfully on average; 11 of the 14 reviews found no difference at all.
Savović J, Turner RM, Mawdsley D, Jones HE, Beynon R, Higgins JPT, Sterne JAC. Association Between Risk-of-Bias Assessments and Results of Randomized Trials in Cochrane Reviews: The ROBES Meta-Epidemiologic Study. American Journal of Epidemiology, 2018 (Meta-analysis | Human)DOI 10.1093/aje/kwx344 In 2,443 randomised trials drawn from 228 meta-analyses, treatment effects were exaggerated by roughly 8 to 13 per cent on average where the risk of bias for random allocation, allocation concealment and blinding was unclear or high.
Perel P, Roberts I, Sena E, Wheble P, Briscoe C, Sandercock P, Macleod M, Mignini LE, Jayaram P, Khan KS. Comparison of treatment effects between animal experiments and clinical trials: systematic review. BMJ, 2007 (Systematic review | Multiple species)DOI 10.1136/bmj.39048.407928.BE For six treatments with a clear clinical result, animal experiment and clinical trial disagreed on several counts: tirilazad shrank the infarct area by 29 per cent in the animal model, while stroke patients given it fared worse.
Hirst JA, Howick J, Aronson JK, Roberts N, Perera R, Koshiaris C, Heneghan C. The need for randomization in animal trials: an overview of systematic reviews. PLoS ONE, 2014 (Systematic review | Multiple species)DOI 10.1371/journal.pone.0098856 Across 31 systematic reviews of animal studies, 29 per cent of the individual studies reported a randomisation, 15 per cent a concealed allocation and 35 per cent a blinded outcome assessment; missing randomisation enlarged the effect estimates.
Bello S, Krogsbøll LT, Gruber J, Zhao ZJ, Fischer D, Hróbjartsson A. Lack of blinding of outcome assessors in animal model experiments implies risk of observer bias. Journal of Clinical Epidemiology, 2014 (Meta-analysis | Multiple species)DOI 10.1016/j.jclinepi.2014.04.008 In ten animal experiments involving 2,450 animals assessed both blinded and unblinded, the unblinded assessment exaggerated the effect by 59 per cent on average; the outcomes were mostly subjective. The figure is carried by three pesticide experiments with a very large deviation: in the remaining seven experiments the exaggeration came to about 18 per cent and was not statistically significant.
Conzemius MG, Evans RB. Caregiver placebo effect for dogs with lameness from osteoarthritis. Journal of the American Veterinary Medical Association, 2012 (Randomised trial | Multiple species)DOI 10.2460/javma.241.10.1314 In the placebo group (58 dogs), owners saw an improvement in 39.7 per cent of cases and veterinary surgeons in 44.8 per cent, while the force plate showed no change in 46 of the 58 dogs.
Gruen ME, Dorman DC, Lascelles BDX. Caregiver placebo effect in analgesic clinical trials for cats with naturally occurring degenerative joint disease-associated pain. Veterinary Record, 2017 (Other | Multiple species)DOI 10.1136/vr.104168 Secondary analysis of the placebo groups: a literature review found six placebo-controlled studies, five of them with usable data on 96 cats; 54 to 74 per cent counted as a success in the owners’ judgement, but only 10 to 63 per cent by objectively measured activity.
Munro BA, Bergen P, Pang DSJ. Randomization, blinding, data handling and sample size estimation in papers published in Veterinary Anaesthesia and Analgesia in 2009 and 2019. Veterinary Anaesthesia and Analgesia, 2022 (Cross-sectional study | Multiple species)DOI 10.1016/j.vaa.2021.09.004 Full description of the randomisation rose between 2009 and 2019 from 13.6 to 85.7 per cent and sample size calculation from 0 to 20 per cent, while blinding stayed at around 50 per cent; no item reached the target that had been set.
Vogt L, Reichlin TS, Nathues C, Würbel H. Authorization of Animal Experiments Is Based on Confidence Rather than Evidence of Scientific Rigor. PLoS Biology, 2016 (Cross-sectional study | Multiple species)DOI 10.1371/journal.pbio.2000598 In 1,277 Swiss animal experiment applications, seven standard measures against bias were described at all in only 2.4 to 19 per cent of the applications, and in 0 to 34 per cent of the publications that came out of them.
Tang X, Shi X, Zhao H, Lu L, Chen Z, Feng Y, Liu L, Duan R, Zhang P, Xu Y, Cui S, Gong F, Fei J, Xu NG, Jing X, Guyatt G, Zhang YQ. Characteristics and quality of clinical practice guidelines addressing acupuncture interventions: a systematic survey of 133 guidelines and 433 acupuncture recommendations. BMJ Open, 2022 (Cross-sectional study | Human)DOI 10.1136/bmjopen-2021-058834 Of 303 recommendations rated with GRADE, 131 were worded strongly, and 104 of those rested on low or very low certainty of the evidence.
Pappachan JM, Lahart IM, Viswanath AK, Borumandi F, Sodi R, Metzendorf MI, Bongaerts B. Parathyroidectomy for adults with primary hyperparathyroidism. Cochrane Database of Systematic Reviews, 2023 (Systematic review | Human)DOI 10.1002/14651858.CD013035.pub2 In the same review (8 randomised trials, 447 adults), certainty was rated moderate for the outcome of cure and very low for bone density at the lumbar spine.
Goehring L, Dorman DC, Osterrieder K, Burgess BA, Dougherty K, Gross P, Neinast C, Pusterla N, Soboll-Hussey G, Lunn DP. Pharmacologic interventions for the treatment of equine herpesvirus-1 in domesticated horses: A systematic review. Journal of Veterinary Internal Medicine, 2024 (Systematic review | Horse)DOI 10.1111/jvim.17016 Of 7,009 publications screened, nine met the criteria; the risk of bias was moderate to high and the groups were small, and the authors saw only a minimal or limited benefit for all the drugs examined, both preventively and after infection.
Hyytiäinen HK, Boström A, Asplund K, Bergh A. A Systematic Review of Complementary and Alternative Veterinary Medicine in Sport and Companion Animals: Electrotherapy. Animals, 2022 (Systematic review | Multiple species)DOI 10.3390/ani13010064 Of 5,385 publications screened, 41 were included, on average 3.7 per procedure; 20 of them carried a high risk of bias and the evidence was sufficient for no clinical application in horses, dogs or cats.
Lai NM, Chang SMW, Ng SS, Tan SL, Chaiyakunapruk N, Stanaway F. Animal-assisted therapy for dementia. Cochrane Database of Systematic Reviews, 2019 (Systematic review | Human)DOI 10.1002/14651858.CD013243.pub2 Nine randomised trials with 305 participants (one of them with horses) yielded low certainty for a slight reduction in depressive symptoms and no clear effect on quality of life; every trial was at high risk of performance bias.
Omar MI, Pal RP, Kelly BD, Bruins HM, Yuan Y, Diemer T, Krausz C, Tournaye H, Kopa Z, Jungwirth A, Minhas S. Benefits of Empiric Nutritional and Medical Therapy for Semen Parameters and Pregnancy and Live Birth Rates in Couples with Idiopathic Infertility: A Systematic Review and Meta-analysis. European Urology, 2019 (Meta-analysis | Human)DOI 10.1016/j.eururo.2018.12.022 From 5,663 records screened, 61 studies were included, 59 of them randomised; under GRADE the certainty of the evidence stayed very low for every outcome.
Ioannidis JPA. The Mass Production of Redundant, Misleading, and Conflicted Systematic Reviews and Meta-analyses. The Milbank Quarterly, 2016 (Other | Species-independent)DOI 10.1111/1468-0009.12210 Between 1991 and 2014 the number of systematic reviews published each year rose by 2,728 per cent, against 153 per cent for all indexed papers; on antidepressants alone, 185 meta-analyses appeared between 2007 and 2014, and the author holds the majority of this output to be redundant or misleading.
Shakir S, Lane S. From a pyramid to an amphitheatre (triangle to circle): embracing the totality of evidence for the postmarketing evaluation of the safety of therapeutic interventions. BMJ Open, 2025 (Other | Human)DOI 10.1136/bmjopen-2025-103538 Argues that the rigid evidence pyramid is unsuited to safety questions after authorisation, and proposes a hierarchy-free model in which randomised and real-world care data complement each other.
ForschungPferd (2026). Understanding evidence levels: what each study design shows and what it does not. ForschungPferd, English. https://forschungpferd.ch/en/evidence/understanding-evidence-levels/