AdvancedReporting guideline: PRISMA 2020Independently reviewed, not yet spot-checked by a human

Systematic review and meta-analysis

Why a systematic review and a meta-analysis are not the same thing, what each box of the PRISMA flow diagram is defending against, why "the search returned nothing" and "nothing survived screening" must never be written as one sentence, and how a result that exists only inside a figure can push a whole review's conclusion off course.

Three words that get used interchangeably, and should not be

NameWhat it doesStatistical pooling?
Narrative reviewAn expert discusses the literature they consider importantNo
Systematic review (SR)The search strategy and the inclusion and exclusion criteria are fixed in advance, so that every qualifying study is found and appraised reproduciblyNot necessarily
Meta-analysis (MA)Effect estimates from several studies are combined statistically into one estimateYes

The relationship between them is that the SR is the process and the MA is one step inside it. An SR without an MA is common and often correct — when the studies differ so much that a pooled number would mean nothing, not pooling is the right decision. An MA without an SR, on the other hand, is a red flag: if you cannot say how these studies were found, there is no answering the question of what population the pooled number estimates.

The PRISMA flow diagram: what each box is defending against

The flow diagram looks like paperwork, but the count at every level answers a “did you miss something, or quietly discard something” question:

StageThe countWhat it defends against
IdentificationHits per database“We searched PubMed, therefore the search was systematic”
Records after deduplicationInflating the total by counting duplicates
ScreeningRecords excluded on title and abstractAn opaque screening process
EligibilityFull texts excluded, with the reason for eachThe single most informative box on the diagram — the distribution of exclusion reasons reveals whether the inclusion criteria were adjusted after the fact
IncludedStudies finally includedRead against every level above, it shows whether the attrition is plausible

Pay particular attention to the exclusion reasons. If an unusual number of full texts were excluded for “does not meet the outcome definition”, the reasonable suspicion is that the outcome definition was narrowed after the results had been seen.

An illustrative PRISMA 2020 flow diagram in two columns. The left column is how many records survive each stage; the four stages, named on a vertical band at the far left, run from Identification through Screening and Eligibility to Included. The right column is where the records leave, each box reached by a horizontal arrow from the left column. The first box lists the hits from four databases (PubMed 812; Embase 936; Cochrane CENTRAL 271; Web of Science 388) for 2407 records identified in total, with 861 duplicates removed shown beside it; the second is the 1546 records screened on title and abstract, with 1394 excluded beside it; the third is the 152 reports sought for retrieval, with 7 not retrieved beside it; the fourth is the 145 reports assessed for eligibility, beside which a taller box itemises the 124 exclusions reason by reason (Not randomised 46; Wrong population 31; Outcome not reported in a usable form 24; Follow-up shorter than 6 months 12; Duplicate report of an included trial 8; Full text not obtainable 3); the bottom box gives the 21 studies included in the review and, separately, the 18 pooled in the meta-analysis, the other 3 having reported no usable variance. Every number in the figure is constructed for teaching and is not the record of any published review.
The four stages of a PRISMA 2020 flow diagram. The left column is what survives, the right column is what leaves, and the two have to reconcile. Every count shown is a constructed teaching example, not data from any published systematic review.Plotting script figures/scripts/D7-sr-figures.R

The box worth stopping on is the one at the bottom right. In this illustration 145 full texts were assessed and 124 were excluded, and the reasons distribute as: not randomised 46, wrong population 31, and outcome not reported in a usable form 24. The signal described just above is whether that third reason swells.

Each count in the left column, minus the box beside it on the right, has to equal the next count down. 2407 records less 861 duplicates leaves 1546; that less the 1394 excluded on title and abstract leaves 152. Reading someone else’s diagram, those subtractions are worth actually doing — where they fail to reconcile, a level has gone unreported.

The bottom box splits into two numbers deliberately: 21 studies entered the review, only 18 entered the pooled analysis, and the 3 in between reported no variance that could be combined. A review count and a pooling count that differ are normal; writing them as one number is not.

Search strategy: reproducibility is the floor, not the ceiling

The search section of an SR should let someone else follow it and arrive at a comparable result. That means stating:

  • Which databases were searched (PubMed / Embase / Cochrane CENTRAL / Web of Science / any national-language databases)
  • The complete search string — usually in a supplement — including how MeSH terms and free-text terms were combined
  • The search date and any date restrictions
  • Language restrictions, if any, stated as such, with the awareness that they introduce language bias
  • Whether anything was added by hand: citation chasing, the reference lists of included studies, asking domain experts

Screening: two people independently, and report the agreement

The standard approach is two reviewers screening independently, with a third adjudicating disagreements. What gets reported is the agreement at the screening stage, usually Cohen’s kappa.

If your SR was screened by one person, or with AI assistance, that belongs in the limitations section of the paper, not only in your working notes. The miss rate of a single reviewer is substantial in the published literature, and readers are entitled to know.

Risk of bias: not a “quality score”

Every included study is assessed for risk of bias, one at a time. The tool depends on the design:

  • RCTs → RoB 2 (five domains: randomisation process, deviations from intended interventions, missing outcome data, outcome measurement, selective reporting)
  • Observational studies → ROBINS-I (which adds confounding and participant selection)
An illustrative RoB 2 traffic-light plot. Each row is one study, named Study A through Study E; each column is one of the five RoB 2 domains, labelled D1 to D5 with the full names given in the lower half of the figure (D1 Randomisation process; D2 Deviations from intended interventions; D3 Missing outcome data; D4 Measurement of the outcome; D5 Selection of the reported result), and one further column, separated by a dashed rule, holds the overall judgement. Every cell is a circle carrying both a colour and a symbol: a green circle with a plus sign for low risk of bias, an amber circle with a question mark for some concerns, a red circle with a minus sign for high risk of bias, so the plot can be read without relying on colour. Read across, the judgements are Study A, D1 to D5 low risk of bias, low risk of bias, low risk of bias, low risk of bias, some concerns, overall some concerns; Study B, D1 to D5 low risk of bias, some concerns, low risk of bias, low risk of bias, low risk of bias, overall some concerns; Study C, D1 to D5 high risk of bias, low risk of bias, some concerns, low risk of bias, low risk of bias, overall high risk of bias; Study D, D1 to D5 low risk of bias, low risk of bias, low risk of bias, low risk of bias, low risk of bias, overall low risk of bias; Study E, D1 to D5 some concerns, high risk of bias, high risk of bias, some concerns, low risk of bias, overall high risk of bias. Of the 25 domain cells, 17 are low risk, 5 are some concerns and 3 are high risk, while the overall column is 1 low, 2 some concerns and 2 high. Below the grid come the legend for the three judgements, the full domain names, and the rule for the overall column: any domain judged high makes the overall high, otherwise a single some-concerns judgement makes the overall some concerns. Every judgement shown is constructed for teaching; no trial was appraised to produce it.
The five RoB 2 domains and the overall judgement, one row per study. Colour and symbol carry the same information, so neither alone is load-bearing. Every judgement shown is a constructed teaching example, not the appraisal of any real trial.Plotting script figures/scripts/D7-sr-figures.R

Read this one from the rightmost column backwards. Study D is green across all five domains, and only then is the overall judgement low risk of bias. Study A has one amber cell, in selective reporting, and the overall is already some concerns. Study E has two red cells and the overall is high. The overall column is not an average of the other five; it is a rule, and any domain judged high makes the overall high. One red cell cannot be pulled back, which is what giving no total score looks like in practice.

Of the 25 cells in this grid, 17 are green, which reads as reassuring; but only 1 of the five studies is low risk of bias overall, and 2 are high. That gap is the whole point of judging domain by domain — a tally of green cells does not offset a red one.

The purpose of a risk of bias assessment is not to rank the studies. It is to decide two things: whether to run a sensitivity analysis that pools only the low risk of bias studies, and whether the final certainty of evidence should be downgraded.

Pooling: when to, and when not to

The technical detail sits on the method pages — forest plots, fixed versus random effects, I² and τ², funnel plots. This section is only about deciding whether to pool at all.

The criterion is clinical and methodological comparability, not statistical feasibility:

  • Populations too different (children and the elderly; treatment-naive and relapsed) → the average applies to nobody
  • Interventions not comparable (a threefold difference in dose, six months’ difference in duration)
  • The outcome is defined differently (“heart failure hospitalisation” can mean entirely different things across studies)
  • Follow-up lengths differ widely

Take the BCG vaccine data used on this site’s method pages: pooling 13 trials gives an RR of 0.49 (95% CI 0.34–0.70), but I² reaches 92.2%. Heterogeneity that high is not there to be suppressed; it is there to be traced. The candidate explanation this dataset is famous for is the latitude of the trial site. Keep in mind what kind of finding that is: a study-level, ecological association. Latitude covaries with the era of the trial, the nutritional status of the population, the BCG strain, the diagnostic criteria and the quality of the trial, and a meta-regression cannot separate them; residual heterogeneity also remains once latitude is in the model. Latitude is therefore a hypothesis-generating candidate, not a settled answer for these data (the heterogeneity page takes those limitations apart one at a time). Even so, tracing the source of the heterogeneity, whether by meta-regression or subgroup analysis, is worth far more than reporting one pooled number.

Certainty of evidence: GRADE

GRADE rates the certainty of evidence for each outcome as high, moderate, low or very low. The starting point depends on the design — RCTs start high, observational studies start low — and there are five reasons to downgrade:

  1. Risk of bias in the included studies
  2. Inconsistency (heterogeneity that cannot be explained)
  3. Indirectness (population, intervention or outcome not quite matching the question)
  4. Imprecision (a confidence interval wide enough to span clinically opposite decisions)
  5. Publication bias

Observational studies have three reasons to upgrade: a very large effect, a dose-response relationship, and plausible confounding that would shrink the observed effect rather than create it.

An illustrative GRADE Summary of Findings table, drawn as a figure with six columns. The first column is the outcome and its follow-up; the second is the number of studies and participants; the third and fourth together form the anticipated absolute effects under one spanning header, giving the events per 1000 people with usual care and the events per 1000 with the intervention together with its confidence interval; the fifth is the relative effect as a risk ratio with its confidence interval; the sixth is the certainty of the evidence, shown as four circles of which the filled ones give the level, with the level name below them and the downgrade codes in brackets below that. The rows read All-cause mortality at 12 months, 12 studies and 4218 participants, 180 per 1000 with usual care against 144 per 1000 (126 to 167) with the intervention, RR 0.80 (0.70 to 0.93), certainty high; Hospital readmission at 6 months, 15 studies and 5104 participants, 320 per 1000 with usual care against 275 per 1000 (243 to 320) with the intervention, RR 0.86 (0.76 to 1.00), certainty moderate with code a; Serious adverse events at 12 months, 9 studies and 2960 participants, 65 per 1000 with usual care against 78 per 1000 (57 to 107) with the intervention, RR 1.20 (0.88 to 1.65), certainty low with code b and c. The last row, Health-related quality of life at 12 months, is a continuous outcome from 6 studies and 1487 participants, so it is reported as a mean difference of 3.4 points (-0.4 to 7.2) on a scale of 0-100 points, higher is better, its relative-effect cell reads not applicable, and its certainty is very low with codes b, c, d. Beneath the table a line notes that the two effect columns are the same estimate written twice, followed by the four downgrade footnotes (a, inconsistency; b, risk of bias; c, imprecision; d, indirectness). Every number and every rating in the table is constructed for teaching and is not the synthesis of any real body of evidence.
The column structure of a GRADE Summary of Findings table: two absolute-effect columns, one relative-effect column, and a certainty column carrying the downgrade codes. Every number and rating shown is a constructed teaching example, not a result from any real body of evidence.Plotting script figures/scripts/D7-sr-figures.R

The table looks like tidy arithmetic, but its two effect columns are the same estimate written twice. Take the first row: 180 events per 1000 with usual care, a relative effect of RR 0.80, and the 144 in the intervention column is those two multiplied together, while the 126 to 167 in brackets is each end of the confidence interval multiplied the same way. The absolute effect is not a second finding — it is the same relative effect applied to one assumed baseline risk, which is why the entire column moves when the baseline population changes.

The second row is where the wording most often goes wrong. Readmission has an RR of 0.86, but the upper limit of the 95% CI is 1.00, sitting exactly on no effect, so the upper limit of the converted intervention risk comes back to the control group’s own 320 per 1000. That row can only be written as not statistically significant, or as evidence that did not detect a difference. It cannot be written as the two groups being the same.

The letters in brackets in the last column are what the table is really for: they are the downgrade codes, spelled out in the footnotes beneath the table, and the rating on its own is only the conclusion while the codes say why. Here the readmission row carries code a for inconsistency and so drops one level from high to moderate; the quality-of-life row carries b, c, d, for risk of bias, imprecision and indirectness, and falls three levels to very low.

The thing most worth remembering from this chapter: some results exist only inside figures

While doing an SR you will use search strings, you will use full-text search, and you may well have a language model help you read. All of those share one blind spot:

A paper’s key result may be drawn inside the plotting area of a figure and never mentioned in the text or the abstract.

This is not hypothetical. In a heart failure study by this site’s author, the plan had been to claim that no study had yet examined a particular interaction. Pulling the most similar papers up by hand revealed that in one of them the interaction p-value appeared only in the rightmost column of Figure 4 — invisible to text search and to full-text grep alike — and the novelty claim had to be narrowed.

The same thing happened once already in this site’s randomised controlled trial chapter: that trial’s CONSORT screening counts appear nowhere in the full text of the Methods or the Results, only inside the Figure 1 image. The PRISMA diagram above is the other face of the same problem — a flow diagram is a convention for putting numbers only inside a picture, and it happens to be the part of a review that most needs checking box by box.

Common misuses

MisuseWhy it is wrong
Calling a one-database search systematicDatabases index different things, and the miss rate from a single source is high
Writing “the search returned nothing” and “nothing survived screening” as one statementThey mean completely different things to a reader, and the first usually means a broken search string
An absolute claim with no search boundaries attached“No study to date has examined X” cannot be checked if you do not say what was searched
Screening with a single reviewer and not disclosing itThe miss rate is substantial and readers are entitled to know
Replacing domain-by-domain assessment with a summed quality score such as JadadIt assumes different biases can offset one another
Forcing a pooled number out of highly heterogeneous studiesTrace the source of the heterogeneity, or do not pool
Treating a GRADE rating as the rating of the whole paperGRADE rates each outcome
Testing funnel plot asymmetry with fewer than 10 studiesSo underpowered as to be uninterpretable (the conventional threshold from the Cochrane Handbook, not a result from these data)
Claiming novelty without having looked at the figures of the most similar studiesThe key result may exist only inside the plotting area
Reading a non-significant pooled result as “the treatment does not work”All you can say is that the available evidence did not detect a difference

Read the figure

The answer comes from the same statistical output that produced this page's figures, not from a number typed in beside them.

The bottom box of the PRISMA diagram deliberately carries two numbers: how many studies the review includes and how many entered the pooled analysis. Why split them?

Show the answer and why

Correct answer: Because 3 studies met the criteria but reported no variance that could be pooled - inclusion and pooling are different things

A study included in a review does not necessarily reach the pooled analysis, and the commonest reason is that it reported no variance that could be combined. Writing both as one number leaves the reader unable to tell which set of studies the pooled estimate represents. The 145 is how many were assessed at full text and the 124 is how many were excluded at that stage; subtracting one from the other gives the number included. Both belong to the layer above and answer whether the screening was transparent, not who was pooled. Note also that the exclusion-reason box repays attention: if unusually many studies were excluded because the outcome was not reported in a usable form, a reasonable suspicion is that the outcome definition was narrowed after the results were seen.

The left column of a PRISMA diagram is what remains at each stage and the right column is what was lost there. What should a reader do with the two columns?

Show the answer and why

Correct answer: Actually do the subtraction. The 1546 is what should remain once duplicates come off the total hits, and if it does not reconcile, a layer has gone unexplained

Each box on the left minus the box beside it on the right has to equal the next box down, and that is the one property of the diagram an outsider can check. 2407 records less 861 duplicates leaves 1546; subtracting the title-and-abstract exclusions from that gives the number sought at full text. These subtractions are worth actually doing, because when they fail to reconcile a layer has gone unexplained, and that layer is usually the least transparent part of the screening. The third answer conflates judgement with bookkeeping: the call in each box is indeed the screeners', but the arithmetic between boxes is not a judgement and can be verified.

In the demonstration RoB 2 figure, seventeen of the twenty-five domain cells are green, yet only one of the five studies is judged low risk overall. Where does that gap come from?

Show the answer and why

Correct answer: Because overall is not an average of the five cells. Only a study green in all five domains counts as low risk, and just 1 of the five is

Overall is a rule, not an average: any domain judged high makes the overall high, and with no high but some concerns the overall is some concerns. So the study green in all five domains is the one that counts as low risk, and a single amber cell is enough to pull an overall to some concerns. Seventeen green cells look like a lot, but they are spread across five studies and cannot offset any red or amber cell. This is what refusing to give a total score looks like in practice, and it is why older instruments in the Jadad mould that add different kinds of bias into one score are no longer recommended: adding them assumes good blinding can compensate for poor randomisation, which it cannot. The point of assessing risk of bias is also not to rank studies but to decide whether to run a sensitivity analysis and whether to downgrade the certainty of the evidence.

Pooling the thirteen BCG vaccine trials leaves a heterogeneity statistic above ninety per cent. What should be done with it?

Show the answer and why

Correct answer: Treat it as something to explain. A value of 92.22 says the trials disagree a great deal, and the response is to look for where the difference comes from

92.22 is not a number to suppress; it is evidence that these studies disagree. Switching to a fixed-effect model does not remove the disagreement, only hides it inside a narrower interval, because a fixed-effect model assumes every study estimates the same quantity and these data are saying that assumption fails. The 0.49 and the upper bound of 0.70 describe the pooled estimate itself, and an interval clear of the null does not make that average applicable to anyone: when populations, interventions, outcome definitions and follow-up differ this much, the average may fit no clinical situation at all. The explanation most often demonstrated on this dataset is the latitude of the trial sites, but that is a study-level association, latitude co-varies with era, nutrition, strain and diagnostic criteria, meta-regression cannot separate them, and residual heterogeneity persists once latitude is in the model. Even so, hunting for the source is worth more than reporting a single pooled number.

In the demonstration Summary of Findings table, the readmission row has a risk ratio below the null with an upper bound sitting exactly on it. How should that row be written?

Show the answer and why

Correct answer: As not statistically significant, or as evidence that detected no difference. The upper bound is 1.00, which returns to the control group's own level

With an upper bound of 1.00 sitting on the null, this row can only be written as not statistically significant, or as evidence that detected no difference. The first answer treats the point estimate as the conclusion and ignores that the interval is saying these data are compatible with no effect. The third answer errs in the opposite direction and more seriously: failing to reach significance is not the same as equivalence, which needs a margin set in advance, and the span from 0.76 to 1.00 holds both an appreciable benefit and none at all. Note also that the two absolute-effect columns of this table are not separate findings; they are the same relative effect applied to an assumed baseline risk, so changing the baseline population changes the whole column.

In one systematic review the all-cause mortality row is rated high certainty and the quality-of-life row very low. What does that mean?

Show the answer and why

Correct answer: That GRADE rates each outcome, not the paper. The 6 studies behind the quality-of-life row are rated separately from the mortality row

GRADE rates every outcome separately, starting from a level set by the design and then downgrading for risk of bias, inconsistency, indirectness, imprecision and publication bias. A review whose primary outcome is high certainty and one of whose secondary outcomes is very low is therefore entirely ordinary, and the second is often the one a headline picks up. Neither 12 studies nor 1487 participants decides a rating: many studies all at serious risk of bias are still downgraded, and many participants whose interval still spans appreciable benefit and appreciable harm are still downgraded for imprecision. What repays reading in this table is the footnote marks in the rightmost column, because the rating is only the conclusion and the marks say why.

Watch next

Narrative vs systematic vs scoping review
ENResearch Masterminds· 9 minSeparate the three kinds of review before anything else — that is exactly what section one below does.
How To Conduct A Systematic Review and Write-Up in 7 Steps (PRISMA, PICO)
ENDr Amina Yonis· 18 minWalks the whole process from PICO to PRISMA. Worth watching before you actually start one.
文獻搜尋 – Cochrane Library + PubMed
繁中Cochrane Taiwan· 96 minIn Traditional Chinese. A hands-on demonstration of turning a PICO question into a search string — the practical companion to section three.
如何選擇一個統合分析的研究題目
繁中杜裕康老師研究室· 11 minIn Traditional Chinese. On choosing a question that can actually be answered — get this step wrong and no amount of elegant statistics later will rescue it.
How to Critically Appraise a Systematic Review: Part 1
ENTerry Shaneyfelt· 8 minAppraising a systematic review as a reader. Read it against the "common misuses" table at the end of this chapter.

Sources and licences

Report a content problem

The statistics on this site are written by AI and reviewed by AI; a human only spot-checks. What you can see may be what we cannot.

The more specific, the more fixable — e.g. which sentence disagrees with which textbook or paper.

Needed only if you want a reply; reports without it are still read.

Sent along with your report

These are attached automatically. You can drop any of them.