Systematic review and meta-analysis
Why a systematic review and a meta-analysis are not the same thing, what each box of the PRISMA flow diagram is defending against, why "the search returned nothing" and "nothing survived screening" must never be written as one sentence, and how a result that exists only inside a figure can push a whole review's conclusion off course.
Three words that get used interchangeably, and should not be
| Name | What it does | Statistical pooling? |
|---|---|---|
| Narrative review | An expert discusses the literature they consider important | No |
| Systematic review (SR) | The search strategy and the inclusion and exclusion criteria are fixed in advance, so that every qualifying study is found and appraised reproducibly | Not necessarily |
| Meta-analysis (MA) | Effect estimates from several studies are combined statistically into one estimate | Yes |
The relationship between them is that the SR is the process and the MA is one step inside it. An SR without an MA is common and often correct — when the studies differ so much that a pooled number would mean nothing, not pooling is the right decision. An MA without an SR, on the other hand, is a red flag: if you cannot say how these studies were found, there is no answering the question of what population the pooled number estimates.
The PRISMA flow diagram: what each box is defending against
The flow diagram looks like paperwork, but the count at every level answers a “did you miss something, or quietly discard something” question:
| Stage | The count | What it defends against |
|---|---|---|
| Identification | Hits per database | “We searched PubMed, therefore the search was systematic” |
| Records after deduplication | Inflating the total by counting duplicates | |
| Screening | Records excluded on title and abstract | An opaque screening process |
| Eligibility | Full texts excluded, with the reason for each | The single most informative box on the diagram — the distribution of exclusion reasons reveals whether the inclusion criteria were adjusted after the fact |
| Included | Studies finally included | Read against every level above, it shows whether the attrition is plausible |
Pay particular attention to the exclusion reasons. If an unusual number of full texts were excluded for “does not meet the outcome definition”, the reasonable suspicion is that the outcome definition was narrowed after the results had been seen.
figures/scripts/D7-sr-figures.RThe box worth stopping on is the one at the bottom right. In this illustration 145 full texts were assessed and 124 were excluded, and the reasons distribute as: not randomised 46, wrong population 31, and outcome not reported in a usable form 24. The signal described just above is whether that third reason swells.
Each count in the left column, minus the box beside it on the right, has to equal the next count down. 2407 records less 861 duplicates leaves 1546; that less the 1394 excluded on title and abstract leaves 152. Reading someone else’s diagram, those subtractions are worth actually doing — where they fail to reconcile, a level has gone unreported.
The bottom box splits into two numbers deliberately: 21 studies entered the review, only 18 entered the pooled analysis, and the 3 in between reported no variance that could be combined. A review count and a pooling count that differ are normal; writing them as one number is not.
Search strategy: reproducibility is the floor, not the ceiling
The search section of an SR should let someone else follow it and arrive at a comparable result. That means stating:
- Which databases were searched (PubMed / Embase / Cochrane CENTRAL / Web of Science / any national-language databases)
- The complete search string — usually in a supplement — including how MeSH terms and free-text terms were combined
- The search date and any date restrictions
- Language restrictions, if any, stated as such, with the awareness that they introduce language bias
- Whether anything was added by hand: citation chasing, the reference lists of included studies, asking domain experts
Screening: two people independently, and report the agreement
The standard approach is two reviewers screening independently, with a third adjudicating disagreements. What gets reported is the agreement at the screening stage, usually Cohen’s kappa.
If your SR was screened by one person, or with AI assistance, that belongs in the limitations section of the paper, not only in your working notes. The miss rate of a single reviewer is substantial in the published literature, and readers are entitled to know.
Risk of bias: not a “quality score”
Every included study is assessed for risk of bias, one at a time. The tool depends on the design:
- RCTs → RoB 2 (five domains: randomisation process, deviations from intended interventions, missing outcome data, outcome measurement, selective reporting)
- Observational studies → ROBINS-I (which adds confounding and participant selection)
figures/scripts/D7-sr-figures.RRead this one from the rightmost column backwards. Study D is green across all five domains, and only then is the overall judgement low risk of bias. Study A has one amber cell, in selective reporting, and the overall is already some concerns. Study E has two red cells and the overall is high. The overall column is not an average of the other five; it is a rule, and any domain judged high makes the overall high. One red cell cannot be pulled back, which is what giving no total score looks like in practice.
Of the 25 cells in this grid, 17 are green, which reads as reassuring; but only 1 of the five studies is low risk of bias overall, and 2 are high. That gap is the whole point of judging domain by domain — a tally of green cells does not offset a red one.
The purpose of a risk of bias assessment is not to rank the studies. It is to decide two things: whether to run a sensitivity analysis that pools only the low risk of bias studies, and whether the final certainty of evidence should be downgraded.
Pooling: when to, and when not to
The technical detail sits on the method pages — forest plots, fixed versus random effects, I² and τ², funnel plots. This section is only about deciding whether to pool at all.
The criterion is clinical and methodological comparability, not statistical feasibility:
- Populations too different (children and the elderly; treatment-naive and relapsed) → the average applies to nobody
- Interventions not comparable (a threefold difference in dose, six months’ difference in duration)
- The outcome is defined differently (“heart failure hospitalisation” can mean entirely different things across studies)
- Follow-up lengths differ widely
Take the BCG vaccine data used on this site’s method pages: pooling 13 trials gives an RR of 0.49 (95% CI 0.34–0.70), but I² reaches 92.2%. Heterogeneity that high is not there to be suppressed; it is there to be traced. The candidate explanation this dataset is famous for is the latitude of the trial site. Keep in mind what kind of finding that is: a study-level, ecological association. Latitude covaries with the era of the trial, the nutritional status of the population, the BCG strain, the diagnostic criteria and the quality of the trial, and a meta-regression cannot separate them; residual heterogeneity also remains once latitude is in the model. Latitude is therefore a hypothesis-generating candidate, not a settled answer for these data (the heterogeneity page takes those limitations apart one at a time). Even so, tracing the source of the heterogeneity, whether by meta-regression or subgroup analysis, is worth far more than reporting one pooled number.
Certainty of evidence: GRADE
GRADE rates the certainty of evidence for each outcome as high, moderate, low or very low. The starting point depends on the design — RCTs start high, observational studies start low — and there are five reasons to downgrade:
- Risk of bias in the included studies
- Inconsistency (heterogeneity that cannot be explained)
- Indirectness (population, intervention or outcome not quite matching the question)
- Imprecision (a confidence interval wide enough to span clinically opposite decisions)
- Publication bias
Observational studies have three reasons to upgrade: a very large effect, a dose-response relationship, and plausible confounding that would shrink the observed effect rather than create it.
figures/scripts/D7-sr-figures.RThe table looks like tidy arithmetic, but its two effect columns are the same estimate written twice. Take the first row: 180 events per 1000 with usual care, a relative effect of RR 0.80, and the 144 in the intervention column is those two multiplied together, while the 126 to 167 in brackets is each end of the confidence interval multiplied the same way. The absolute effect is not a second finding — it is the same relative effect applied to one assumed baseline risk, which is why the entire column moves when the baseline population changes.
The second row is where the wording most often goes wrong. Readmission has an RR of 0.86, but the upper limit of the 95% CI is 1.00, sitting exactly on no effect, so the upper limit of the converted intervention risk comes back to the control group’s own 320 per 1000. That row can only be written as not statistically significant, or as evidence that did not detect a difference. It cannot be written as the two groups being the same.
The letters in brackets in the last column are what the table is really for: they are the downgrade codes, spelled out in the footnotes beneath the table, and the rating on its own is only the conclusion while the codes say why. Here the readmission row carries code a for inconsistency and so drops one level from high to moderate; the quality-of-life row carries b, c, d, for risk of bias, imprecision and indirectness, and falls three levels to very low.
The thing most worth remembering from this chapter: some results exist only inside figures
While doing an SR you will use search strings, you will use full-text search, and you may well have a language model help you read. All of those share one blind spot:
A paper’s key result may be drawn inside the plotting area of a figure and never mentioned in the text or the abstract.
This is not hypothetical. In a heart failure study by this site’s author, the plan had been to claim that no study had yet examined a particular interaction. Pulling the most similar papers up by hand revealed that in one of them the interaction p-value appeared only in the rightmost column of Figure 4 — invisible to text search and to full-text grep alike — and the novelty claim had to be narrowed.
The same thing happened once already in this site’s randomised controlled trial chapter: that trial’s CONSORT screening counts appear nowhere in the full text of the Methods or the Results, only inside the Figure 1 image. The PRISMA diagram above is the other face of the same problem — a flow diagram is a convention for putting numbers only inside a picture, and it happens to be the part of a review that most needs checking box by box.
Common misuses
| Misuse | Why it is wrong |
|---|---|
| Calling a one-database search systematic | Databases index different things, and the miss rate from a single source is high |
| Writing “the search returned nothing” and “nothing survived screening” as one statement | They mean completely different things to a reader, and the first usually means a broken search string |
| An absolute claim with no search boundaries attached | “No study to date has examined X” cannot be checked if you do not say what was searched |
| Screening with a single reviewer and not disclosing it | The miss rate is substantial and readers are entitled to know |
| Replacing domain-by-domain assessment with a summed quality score such as Jadad | It assumes different biases can offset one another |
| Forcing a pooled number out of highly heterogeneous studies | Trace the source of the heterogeneity, or do not pool |
| Treating a GRADE rating as the rating of the whole paper | GRADE rates each outcome |
| Testing funnel plot asymmetry with fewer than 10 studies | So underpowered as to be uninterpretable (the conventional threshold from the Cochrane Handbook, not a result from these data) |
| Claiming novelty without having looked at the figures of the most similar studies | The key result may exist only inside the plotting area |
| Reading a non-significant pooled result as “the treatment does not work” | All you can say is that the available evidence did not detect a difference |
Read the figure
The answer comes from the same statistical output that produced this page's figures, not from a number typed in beside them.
The bottom box of the PRISMA diagram deliberately carries two numbers: how many studies the review includes and how many entered the pooled analysis. Why split them?
Show the answer and why
Correct answer: Because 3 studies met the criteria but reported no variance that could be pooled - inclusion and pooling are different things
A study included in a review does not necessarily reach the pooled analysis, and the commonest reason is that it reported no variance that could be combined. Writing both as one number leaves the reader unable to tell which set of studies the pooled estimate represents. The 145 is how many were assessed at full text and the 124 is how many were excluded at that stage; subtracting one from the other gives the number included. Both belong to the layer above and answer whether the screening was transparent, not who was pooled. Note also that the exclusion-reason box repays attention: if unusually many studies were excluded because the outcome was not reported in a usable form, a reasonable suspicion is that the outcome definition was narrowed after the results were seen.
The left column of a PRISMA diagram is what remains at each stage and the right column is what was lost there. What should a reader do with the two columns?
Show the answer and why
Correct answer: Actually do the subtraction. The 1546 is what should remain once duplicates come off the total hits, and if it does not reconcile, a layer has gone unexplained
Each box on the left minus the box beside it on the right has to equal the next box down, and that is the one property of the diagram an outsider can check. 2407 records less 861 duplicates leaves 1546; subtracting the title-and-abstract exclusions from that gives the number sought at full text. These subtractions are worth actually doing, because when they fail to reconcile a layer has gone unexplained, and that layer is usually the least transparent part of the screening. The third answer conflates judgement with bookkeeping: the call in each box is indeed the screeners', but the arithmetic between boxes is not a judgement and can be verified.
In the demonstration RoB 2 figure, seventeen of the twenty-five domain cells are green, yet only one of the five studies is judged low risk overall. Where does that gap come from?
Show the answer and why
Correct answer: Because overall is not an average of the five cells. Only a study green in all five domains counts as low risk, and just 1 of the five is
Overall is a rule, not an average: any domain judged high makes the overall high, and with no high but some concerns the overall is some concerns. So the study green in all five domains is the one that counts as low risk, and a single amber cell is enough to pull an overall to some concerns. Seventeen green cells look like a lot, but they are spread across five studies and cannot offset any red or amber cell. This is what refusing to give a total score looks like in practice, and it is why older instruments in the Jadad mould that add different kinds of bias into one score are no longer recommended: adding them assumes good blinding can compensate for poor randomisation, which it cannot. The point of assessing risk of bias is also not to rank studies but to decide whether to run a sensitivity analysis and whether to downgrade the certainty of the evidence.
Pooling the thirteen BCG vaccine trials leaves a heterogeneity statistic above ninety per cent. What should be done with it?
Show the answer and why
Correct answer: Treat it as something to explain. A value of 92.22 says the trials disagree a great deal, and the response is to look for where the difference comes from
92.22 is not a number to suppress; it is evidence that these studies disagree. Switching to a fixed-effect model does not remove the disagreement, only hides it inside a narrower interval, because a fixed-effect model assumes every study estimates the same quantity and these data are saying that assumption fails. The 0.49 and the upper bound of 0.70 describe the pooled estimate itself, and an interval clear of the null does not make that average applicable to anyone: when populations, interventions, outcome definitions and follow-up differ this much, the average may fit no clinical situation at all. The explanation most often demonstrated on this dataset is the latitude of the trial sites, but that is a study-level association, latitude co-varies with era, nutrition, strain and diagnostic criteria, meta-regression cannot separate them, and residual heterogeneity persists once latitude is in the model. Even so, hunting for the source is worth more than reporting a single pooled number.
In the demonstration Summary of Findings table, the readmission row has a risk ratio below the null with an upper bound sitting exactly on it. How should that row be written?
Show the answer and why
Correct answer: As not statistically significant, or as evidence that detected no difference. The upper bound is 1.00, which returns to the control group's own level
With an upper bound of 1.00 sitting on the null, this row can only be written as not statistically significant, or as evidence that detected no difference. The first answer treats the point estimate as the conclusion and ignores that the interval is saying these data are compatible with no effect. The third answer errs in the opposite direction and more seriously: failing to reach significance is not the same as equivalence, which needs a margin set in advance, and the span from 0.76 to 1.00 holds both an appreciable benefit and none at all. Note also that the two absolute-effect columns of this table are not separate findings; they are the same relative effect applied to an assumed baseline risk, so changing the baseline population changes the whole column.
In one systematic review the all-cause mortality row is rated high certainty and the quality-of-life row very low. What does that mean?
Show the answer and why
Correct answer: That GRADE rates each outcome, not the paper. The 6 studies behind the quality-of-life row are rated separately from the mortality row
GRADE rates every outcome separately, starting from a level set by the design and then downgrading for risk of bias, inconsistency, indirectness, imprecision and publication bias. A review whose primary outcome is high certainty and one of whose secondary outcomes is very low is therefore entirely ordinary, and the second is often the one a headline picks up. Neither 12 studies nor 1487 participants decides a rating: many studies all at serious risk of bias are still downgraded, and many participants whose interval still spans appreciable benefit and appreciable harm are still downgraded for imprecision. What repays reading in this table is the footnote marks in the rightmost column, because the rating is only the conclusion and the marks say why.
Methods used in this chapter
- Effect measures and their variances
- Fixed-effect and random-effects models
- Heterogeneity — I², τ² and prediction intervals
- Forest and funnel plots
- Network meta-analysis — indirect comparison and treatment ranking
- Cumulative meta-analysis and information size
- Sensitivity and influence analysis
- Meta-analysis of diagnostic test accuracy
- Agreement and the kappa family
- Multiple comparisons and subgroup analyses
Watch next
Narrative vs systematic vs scoping review
How To Conduct A Systematic Review and Write-Up in 7 Steps (PRISMA, PICO)
文獻搜尋 – Cochrane Library + PubMed
如何選擇一個統合分析的研究題目
How to Critically Appraise a Systematic Review: Part 1Sources and licences
- The PRISMA 2020 statement: an updated guideline for reporting systematic reviewsCC BYThe chapter's process skeleton and section order follow the PRISMA 2020 items. The prose is original.