Diagnostic accuracy study
Sensitivity and specificity are properties of a test, but the numbers you get depend on whom you enrolled and whom you verified — this chapter uses live figures to show two ways a good test is made to look better, and why the same test reports a sensitivity three times higher in one ward than another.
What this chapter covers, and what the method pages cover
How sensitivity, specificity, likelihood ratios and ROC curves are calculated and read belongs to the method pages — the B4 family.
This chapter asks the design-level questions instead: which people did these numbers come from, who was left out, and can the reference standard itself be trusted. Because the most awkward thing about diagnostic accuracy is this:
The reference standard: the ground the whole study stands on
A diagnostic accuracy study compares the index test against a reference standard (formerly “gold standard”). Every number it produces means something only if the reference standard is right.
In practice that premise never holds completely:
| Situation | Consequence |
|---|---|
| The reference standard is itself imperfect | The index test’s accuracy is usually underestimated — it was right and got marked wrong. That holds when the two tests’ errors are independent; when they are correlated, the bias can run the other way (see incorporation bias below) |
| The reference standard is invasive (biopsy, angiography) | It cannot be applied to everyone → the verification bias in the next section |
| No single reference standard exists | You need a composite reference standard (consensus from several tests plus clinical follow-up), which may itself contain the index test |
| Whoever reads the reference standard knows the index test result | The reading is contaminated and the agreement between the two is inflated by hand |
Verification bias: what happens when only the positives get verified
The most common design compromise is this: the reference standard is too invasive or too expensive, so it is applied only to those who tested positive. Everyone who tested negative is written down as disease-free. That is verification bias, also called work-up bias.
Its direction can be derived, and you can watch it happen by simulating on one dataset. The table below lowers the proportion of test-negative people who receive the reference standard, from everyone down to one in ten, and reports the apparent sensitivity and specificity each time:
| Proportion of test-negatives verified | Expected number receiving the reference standard | Apparent sensitivity | Apparent specificity |
|---|---|---|---|
| 100% | 113.0 | 0.634 | 0.806 |
| 70% | 91.1 | 0.712 | 0.744 |
| 50% | 76.5 | 0.776 | 0.674 |
| 30% | 61.9 | 0.852 | 0.554 |
| 10% | 47.3 | 0.945 | 0.293 |
(The count column is an expected value, not the result of one draw — the script multiplies the proportion by the number of test-negatives, which is why fractions appear.)
The truth is fixed at sensitivity 0.634 and specificity 0.806. Verify only half the test-negatives and the apparent sensitivity climbs to 0.776 while the specificity falls to 0.674.
That diagram looks like this:
figures/scripts/D4-stard.REach index test result splits into performed and not performed. Of the 214 positives, 205 were verified; of the 590 negatives, only 236 were, leaving 354 in the red box. The leftmost column of the table above — the proportion of test-negatives verified — is the ratio between that box and the one to its left. In this diagram it is 40%, which sits between the 50% and 30% rows of the table.
Notice also that the 2×2 hangs only off the performed boxes — and not even off all of them. Of the 822 who took the index test, 455 received the reference standard, but the 2×2 is built on 441: the 14 verified inconclusive results have no cell to sit in, because an inconclusive index test is neither a positive nor a negative call, and STARD asks for them to be reported separately rather than folded into either column. That is the whole point: the denominator behind the sensitivity and specificity a paper prints is 441, and it is the reader who has to find it and subtract to see who was left outside the table — the 367 never verified, plus the 14 verified but uncountable.
A softer variant is differential verification: test-positives get the invasive reference standard while test-negatives get a weaker one — clinical follow-up, say, or a second imaging study. Nobody is left unverified, so the flow diagram looks complete, but the two groups were not held to the same standard. This is not the same bias as partial verification, and its direction depends on how the weaker standard misclassifies: a follow-up that misses slowly progressing disease will file those patients as true negatives and inflate sensitivity, while one that over-calls will do the reverse. The question that catches it is the one QUADAS-2 asks outright — did every patient receive the same reference standard?
The spectrum effect: one test, a threefold difference in sensitivity
Split the same dataset by disease severity (WFNS grade). Same test, same cut-off:
| Stratum | n | Diseased n | Disease-free n | Prevalence | Sensitivity (95% CI) | Specificity (95% CI) |
|---|---|---|---|---|---|---|
| WFNS 1-2 | 71 | 14 | 57 | 19.7% | 0.286 (0.12–0.55) | 0.947 (0.86–0.98) |
| WFNS 3-5 | 42 | 27 | 15 | 64.3% | 0.815 (0.63–0.92) | 0.267 (0.11–0.52) |
Sensitivity in the milder stratum is 0.286; in the severe stratum it is 0.815 — close to a threefold difference, and the two confidence intervals do not overlap. Specificity moves the other way.
Why would specificity fall? One plausible reading: in the higher-WFNS stratum, even the patients with a good outcome have sustained more brain injury, so their S100β runs higher, so more of them cross the same cut-off and are called false positives — a larger share of the disease-free group is misclassified, and specificity drops. That is a plausible explanation built from the relationship between stratum and marker concentration, not a causal mechanism these data can confirm: confirming it would mean comparing the S100β distribution of the good-outcome patients across the two strata directly, rather than looking only at accuracy indices.
The two extra columns in the table are a reminder of where each denominator comes from: the precision of sensitivity is governed only by the number of diseased patients, and the precision of specificity only by the number of disease-free ones. The severe stratum contains just 15 disease-free patients, which is why its specificity interval is wide enough to carry almost no information — how much of that drop is real and how much is sampling noise is something these data cannot separate.
This is the spectrum effect, and the reason is blunt: severely ill patients have higher biomarker concentrations and are easier to detect, and if the “healthy controls” are people with no symptoms at all, they are easier to call negative.
Sample size: the confidence interval tells you whether you had enough
The table above demonstrates something else. The sensitivity interval in the milder stratum runs from 0.12 to 0.55 — wide enough to be nearly uninformative. That stratum contains only 14 diseased patients, and that number is the denominator of its sensitivity.
Sample size in a diagnostic accuracy study is set by the number of cases and the number of controls, not by the number of people enrolled. A study of 1000 participants containing 20 confirmed cases has a sensitivity whose precision is decided by those 20.
The design choices available
| Design | How it works | Advantage / risk |
|---|---|---|
| Consecutive enrolment | Enrol every eligible patient in order | Closest to clinical reality, full spectrum; slow and expensive |
| Two-gate | Recruit confirmed cases and healthy controls separately | Fast and cheap; accuracy is systematically inflated, see the previous section |
| Nested in a cohort | Sample from within an existing cohort | Sensible spectrum and efficient; limited to whatever the cohort originally collected |
| Randomised diagnostic trial | Randomise patients to different testing strategies and compare clinical outcomes rather than accuracy | The only design that answers “are patients better off when this test is used” |
Blinding
Blinding is needed in three directions:
- Whoever reads the index test does not know the reference standard result
- Whoever reads the reference standard does not know the index test result
- Neither knows the other clinical information — unless that information would be available in practice anyway, in which case say so
For tests requiring human interpretation, such as imaging and pathology, this matters more than anywhere else, and inter-reader agreement has to be reported — usually kappa, which has a property that catches people out: two datasets with the same observed agreement can differ threefold in kappa purely because the abnormal reading is rarer in one of them.
Common misuses
| Misuse | Why it is wrong |
|---|---|
| Applying the reference standard only to test-positives and calling the rest disease-free | Verification bias: sensitivity inflated and specificity deflated at once |
| Defining the reference standard so that it includes the index test | Incorporation bias, systematic inflation |
| Running the study as “confirmed cases versus healthy volunteers” | Two-gate sampling, a polarised spectrum, heavily inflated accuracy |
| Carrying a published sensitivity straight over to your own patients | Spectrum effect: the same test can differ threefold across severity strata |
| Reporting sensitivity without a confidence interval | The denominator is the case count, and it is often small enough that the interval carries no information |
| Treating accuracy as evidence that patients benefit | That needs a trial with clinical outcomes as the endpoint |
| Reading the tests unblinded | The two results contaminate each other and their agreement is inflated by hand |
| Choosing the best cut-off in the same data used to report accuracy | Optimism bias; a cut-off needs independent validation |
Where the numbers on this page come from
Every figure quoted here comes from the analysis scripts behind method pages B4-01 and B4-02:
/opt/homebrew/bin/Rscript figures/scripts/B4-01-sens-spec.R
/opt/homebrew/bin/Rscript figures/scripts/B4-02-ppv-npv.R
/opt/homebrew/bin/Rscript figures/scripts/D4-stard.R
The third script produces only the STARD diagram and its figures/out/D4-stard.stats.json. That file is marked hypothetical: true, because the counts in the diagram were invented to show the structure rather than measured from pROC::aSAH — in that dataset every patient received the reference standard, so it has no unverified branch to draw.
Read the figure
The answer comes from the same statistical output that produced this page's figures, not from a number typed in beside them.
The reference standard is applied only to people who test positive, and those who test negative are taken to be disease free. Which way do the apparent sensitivity and specificity move?
Show the answer and why
Correct answer: Sensitivity is overstated and specificity understated. Verifying half the negatives lifts apparent sensitivity to 0.776
Among the negatives left unverified, true negatives vastly outnumber false negatives, which is simply what it means for a test to have some accuracy. Dropping that whole group strips large numbers of true negatives out of the denominator of specificity, which falls to 0.674, while the denominator of sensitivity loses false negatives, so apparent sensitivity climbs to 0.776. The true sensitivity was 0.634 throughout: the test never changed, only who was counted. The two measures therefore move in opposite directions rather than deteriorating together. The third statement fails on the claim that the denominator of sensitivity is untouched, since false negatives are diseased people who test negative and sit precisely in the group that was skipped. This bias never labels itself in a paper; it has to be inferred from how many people received the reference standard, which is why the STARD diagram exists.
Splitting the same dataset by disease severity, with the same test and the same cut-off, leaves the two strata with sensitivities differing by nearly threefold. What does that show?
Show the answer and why
Correct answer: It shows sensitivity depends on which patients were enrolled. In the severe stratum it is 0.815, because sicker people have higher marker concentrations and cross the same cut-off more often
Sensitivity and specificity do not change with prevalence, which is where the 0.197 answer conflates two things. What they do depend on is who is being tested: 0.286 and 0.815 come from one test at one cut-off, and the only difference is that one stratum is mild and the other severe. Nor is this instability in the measurement, since sicker people have higher marker concentrations and cross the cut-off more often, a predictable direction rather than noise. This is exactly why a two-gate design of confirmed cases against healthy volunteers produces beautiful accuracy that cannot be used: it answers whether the test separates severe disease from health, whereas the clinical question is whether it helps with a patient whose presentation is unclear. Note too that the severe stratum contains barely a dozen people without the disease, so its specificity interval is wide enough to carry almost no information.
The sensitivity interval in the mild stratum is so wide it carries almost no information. Which number sets that width?
Show the answer and why
Correct answer: The 14 people with the disease, because the denominator of sensitivity holds only diseased people
Sensitivity is the proportion of diseased people the test picks up, so its denominator contains only diseased people. The mild stratum has 71 people in it, but only 14 of them have the disease, and those 14 set the precision; the remaining 57 affect specificity alone. A study of a thousand people with twenty confirmed cases therefore has the sensitivity precision of twenty people. That is why a diagnostic accuracy study should report the number of cases and the number of controls rather than the total, since a large total may have nothing to do with the measure you care about.
In the demonstration STARD diagram 822 people had the index test, 455 had the reference standard, and the 2 by 2 table is built on 441. What is the gap between 455 and 441?
Show the answer and why
Correct answer: The 14 people with an inconclusive index test - they did have the reference standard, but inconclusive is neither a positive nor a negative reading
822 minus 455 is 367, the people who never had the reference standard and who left one layer earlier. 455 minus 441 is 14, which is the gap this question asks about: people who did have the reference standard but have no cell to sit in. Both groups end up outside the 2 by 2 table for entirely different reasons, and since the sensitivity and specificity a paper reports have a denominator of 441, the reader has to do the subtraction to see who is missing. The 354 is the unverified box on the negative branch, part of the 367 and the place where verification bias lives. Folding inconclusive results into either the positive or the negative column would let a group that could not be classified at all pull that column's accuracy around.
In the demonstration diagram nearly all 214 test-positive people had the reference standard, while only some of the 590 test-negative ones did. What does that asymmetry mean?
Show the answer and why
Correct answer: It means only 236 people on the negative branch were actually verified, and the rest were assumed disease free rather than confirmed to be
On the positive branch 205 of 214 had the reference standard; on the negative branch only 236 of 590 did. That asymmetry is partial verification: the unverified negatives are written into the table as disease free, and some of them are false negatives. The second statement assumes the positive branch already captures everyone with the disease, but a test that caught everyone would not need an accuracy study. The third compares raw counts: the positive branch misses a single-digit number while the negative branch misses more than three hundred, and the degree of asymmetry is exactly the point. Reading a paper calls for the same arithmetic, subtracting the number who received the reference standard from the number enrolled and then asking how the remainder were handled.
The sensitivity in the main analysis comes with a 95% confidence interval. What does its width tell the reader?
Show the answer and why
Correct answer: The lower bound for sensitivity is only 0.481, so these data cannot rule out missing one of every two diseased people
0.634 is the point estimate and it says nothing about what these data can exclude. The lower bound of 0.481 is the number to read: this dataset is compatible with fewer than half of diseased people being detected. The 0.700 is the lower bound for specificity, whose denominator is the people without the disease, and since this cohort has more of those than of cases, the specificity interval is narrower to begin with, so the two measures are not equally precise. Reporting a sensitivity without its interval presents an estimate resting on a few dozen people as though it were a settled number.
Methods used in this chapter
- The 2x2 table, sensitivity and specificity
- Predictive values and prevalence
- Likelihood ratios and the Fagan nomogram
- ROC curves and the area under them
- Choosing a cut-off, and comparing two ROC curves
- Agreement and the kappa family
- Bland-Altman analysis and the limits of agreement
- Meta-analysis of diagnostic test accuracy
- Selection bias
- Measurement error and misclassification
Watch next
Cohort Studies: A Brief OverviewSources and licences
- STARD 2015: an updated list of essential items for reporting diagnostic accuracy studiesCC BYThe chapter's section order follows the STARD 2015 items. The prose is original.