AdvancedReporting guideline: STARD 2015Independently reviewed, not yet spot-checked by a human

Diagnostic accuracy study

Sensitivity and specificity are properties of a test, but the numbers you get depend on whom you enrolled and whom you verified — this chapter uses live figures to show two ways a good test is made to look better, and why the same test reports a sensitivity three times higher in one ward than another.

What this chapter covers, and what the method pages cover

How sensitivity, specificity, likelihood ratios and ROC curves are calculated and read belongs to the method pages — the B4 family.

This chapter asks the design-level questions instead: which people did these numbers come from, who was left out, and can the reference standard itself be trusted. Because the most awkward thing about diagnostic accuracy is this:

The reference standard: the ground the whole study stands on

A diagnostic accuracy study compares the index test against a reference standard (formerly “gold standard”). Every number it produces means something only if the reference standard is right.

In practice that premise never holds completely:

SituationConsequence
The reference standard is itself imperfectThe index test’s accuracy is usually underestimated — it was right and got marked wrong. That holds when the two tests’ errors are independent; when they are correlated, the bias can run the other way (see incorporation bias below)
The reference standard is invasive (biopsy, angiography)It cannot be applied to everyone → the verification bias in the next section
No single reference standard existsYou need a composite reference standard (consensus from several tests plus clinical follow-up), which may itself contain the index test
Whoever reads the reference standard knows the index test resultThe reading is contaminated and the agreement between the two is inflated by hand

Verification bias: what happens when only the positives get verified

The most common design compromise is this: the reference standard is too invasive or too expensive, so it is applied only to those who tested positive. Everyone who tested negative is written down as disease-free. That is verification bias, also called work-up bias.

Its direction can be derived, and you can watch it happen by simulating on one dataset. The table below lowers the proportion of test-negative people who receive the reference standard, from everyone down to one in ten, and reports the apparent sensitivity and specificity each time:

Proportion of test-negatives verifiedExpected number receiving the reference standardApparent sensitivityApparent specificity
100%113.00.6340.806
70%91.10.7120.744
50%76.50.7760.674
30%61.90.8520.554
10%47.30.9450.293

(The count column is an expected value, not the result of one draw — the script multiplies the proportion by the number of test-negatives, which is why fractions appear.)

The truth is fixed at sensitivity 0.634 and specificity 0.806. Verify only half the test-negatives and the apparent sensitivity climbs to 0.776 while the specificity falls to 0.674.

That diagram looks like this:

A STARD flow diagram in three stages, top to bottom, each on a pale band with its stage name set vertically down the left edge. Enrolment: a box at the top reads eligible patients, 980; on the way down, a branch to the right leads to a box of 158 excluded before the index test, with three reasons and counts beneath it (Did not meet inclusion criteria, 96; Declined to participate, 41; Incomplete baseline records, 21); the trunk continues down into underwent the index test, 822. Middle stage: the index test result splits into three columns, Index test positive, 214; Index test inconclusive, 18; Index test negative, 590. Each column then splits into two boxes side by side, reference standard performed on the left and reference standard NOT performed on the right: 205 against 9 for the positives, 14 against 4 for the inconclusives, and 236 against 354 for the negatives. The not-performed box in the negative column is outlined in red and fed by a red arrow: that is the partial-verification branch. Results: one box per column, hanging only off the performed box, giving the number with the target condition present and the number with it absent. A red panel at the foot states that 354 test-negatives never received the reference standard; that 455 patients did receive it but only 441 of them enter the 2×2, because the 14 verified inconclusive results have no cell to sit in and STARD leaves them out; and that on those 441, sensitivity and specificity come out at 0.884 and 0.853. A final line of red text notes that all the counts are hypothetical.
A STARD flow diagram. The counts inside it are hypothetical, invented to show the structure, and belong to no published study. The part to look at is the right-hand half of the middle stage: every index test result splits into reference standard performed and reference standard not performed, and the red box in the negative column is where partial verification bias lives. A real paper will not outline it in red for you.Plotting script figures/scripts/D4-stard.R

Each index test result splits into performed and not performed. Of the 214 positives, 205 were verified; of the 590 negatives, only 236 were, leaving 354 in the red box. The leftmost column of the table above — the proportion of test-negatives verified — is the ratio between that box and the one to its left. In this diagram it is 40%, which sits between the 50% and 30% rows of the table.

Notice also that the 2×2 hangs only off the performed boxes — and not even off all of them. Of the 822 who took the index test, 455 received the reference standard, but the 2×2 is built on 441: the 14 verified inconclusive results have no cell to sit in, because an inconclusive index test is neither a positive nor a negative call, and STARD asks for them to be reported separately rather than folded into either column. That is the whole point: the denominator behind the sensitivity and specificity a paper prints is 441, and it is the reader who has to find it and subtract to see who was left outside the table — the 367 never verified, plus the 14 verified but uncountable.

A softer variant is differential verification: test-positives get the invasive reference standard while test-negatives get a weaker one — clinical follow-up, say, or a second imaging study. Nobody is left unverified, so the flow diagram looks complete, but the two groups were not held to the same standard. This is not the same bias as partial verification, and its direction depends on how the weaker standard misclassifies: a follow-up that misses slowly progressing disease will file those patients as true negatives and inflate sensitivity, while one that over-calls will do the reverse. The question that catches it is the one QUADAS-2 asks outright — did every patient receive the same reference standard?

The spectrum effect: one test, a threefold difference in sensitivity

Split the same dataset by disease severity (WFNS grade). Same test, same cut-off:

StratumnDiseased nDisease-free nPrevalenceSensitivity (95% CI)Specificity (95% CI)
WFNS 1-271145719.7%0.286 (0.12–0.55)0.947 (0.86–0.98)
WFNS 3-542271564.3%0.815 (0.63–0.92)0.267 (0.11–0.52)

Sensitivity in the milder stratum is 0.286; in the severe stratum it is 0.815 — close to a threefold difference, and the two confidence intervals do not overlap. Specificity moves the other way.

Why would specificity fall? One plausible reading: in the higher-WFNS stratum, even the patients with a good outcome have sustained more brain injury, so their S100β runs higher, so more of them cross the same cut-off and are called false positives — a larger share of the disease-free group is misclassified, and specificity drops. That is a plausible explanation built from the relationship between stratum and marker concentration, not a causal mechanism these data can confirm: confirming it would mean comparing the S100β distribution of the good-outcome patients across the two strata directly, rather than looking only at accuracy indices.

The two extra columns in the table are a reminder of where each denominator comes from: the precision of sensitivity is governed only by the number of diseased patients, and the precision of specificity only by the number of disease-free ones. The severe stratum contains just 15 disease-free patients, which is why its specificity interval is wide enough to carry almost no information — how much of that drop is real and how much is sampling noise is something these data cannot separate.

This is the spectrum effect, and the reason is blunt: severely ill patients have higher biomarker concentrations and are easier to detect, and if the “healthy controls” are people with no symptoms at all, they are easier to call negative.

Sample size: the confidence interval tells you whether you had enough

The table above demonstrates something else. The sensitivity interval in the milder stratum runs from 0.12 to 0.55 — wide enough to be nearly uninformative. That stratum contains only 14 diseased patients, and that number is the denominator of its sensitivity.

Sample size in a diagnostic accuracy study is set by the number of cases and the number of controls, not by the number of people enrolled. A study of 1000 participants containing 20 confirmed cases has a sensitivity whose precision is decided by those 20.

The design choices available

DesignHow it worksAdvantage / risk
Consecutive enrolmentEnrol every eligible patient in orderClosest to clinical reality, full spectrum; slow and expensive
Two-gateRecruit confirmed cases and healthy controls separatelyFast and cheap; accuracy is systematically inflated, see the previous section
Nested in a cohortSample from within an existing cohortSensible spectrum and efficient; limited to whatever the cohort originally collected
Randomised diagnostic trialRandomise patients to different testing strategies and compare clinical outcomes rather than accuracyThe only design that answers “are patients better off when this test is used”

Blinding

Blinding is needed in three directions:

  1. Whoever reads the index test does not know the reference standard result
  2. Whoever reads the reference standard does not know the index test result
  3. Neither knows the other clinical information — unless that information would be available in practice anyway, in which case say so

For tests requiring human interpretation, such as imaging and pathology, this matters more than anywhere else, and inter-reader agreement has to be reported — usually kappa, which has a property that catches people out: two datasets with the same observed agreement can differ threefold in kappa purely because the abnormal reading is rarer in one of them.

Common misuses

MisuseWhy it is wrong
Applying the reference standard only to test-positives and calling the rest disease-freeVerification bias: sensitivity inflated and specificity deflated at once
Defining the reference standard so that it includes the index testIncorporation bias, systematic inflation
Running the study as “confirmed cases versus healthy volunteers”Two-gate sampling, a polarised spectrum, heavily inflated accuracy
Carrying a published sensitivity straight over to your own patientsSpectrum effect: the same test can differ threefold across severity strata
Reporting sensitivity without a confidence intervalThe denominator is the case count, and it is often small enough that the interval carries no information
Treating accuracy as evidence that patients benefitThat needs a trial with clinical outcomes as the endpoint
Reading the tests unblindedThe two results contaminate each other and their agreement is inflated by hand
Choosing the best cut-off in the same data used to report accuracyOptimism bias; a cut-off needs independent validation

Where the numbers on this page come from

Every figure quoted here comes from the analysis scripts behind method pages B4-01 and B4-02:

/opt/homebrew/bin/Rscript figures/scripts/B4-01-sens-spec.R
/opt/homebrew/bin/Rscript figures/scripts/B4-02-ppv-npv.R
/opt/homebrew/bin/Rscript figures/scripts/D4-stard.R

The third script produces only the STARD diagram and its figures/out/D4-stard.stats.json. That file is marked hypothetical: true, because the counts in the diagram were invented to show the structure rather than measured from pROC::aSAH — in that dataset every patient received the reference standard, so it has no unverified branch to draw.

Read the figure

The answer comes from the same statistical output that produced this page's figures, not from a number typed in beside them.

The reference standard is applied only to people who test positive, and those who test negative are taken to be disease free. Which way do the apparent sensitivity and specificity move?

Show the answer and why

Correct answer: Sensitivity is overstated and specificity understated. Verifying half the negatives lifts apparent sensitivity to 0.776

Among the negatives left unverified, true negatives vastly outnumber false negatives, which is simply what it means for a test to have some accuracy. Dropping that whole group strips large numbers of true negatives out of the denominator of specificity, which falls to 0.674, while the denominator of sensitivity loses false negatives, so apparent sensitivity climbs to 0.776. The true sensitivity was 0.634 throughout: the test never changed, only who was counted. The two measures therefore move in opposite directions rather than deteriorating together. The third statement fails on the claim that the denominator of sensitivity is untouched, since false negatives are diseased people who test negative and sit precisely in the group that was skipped. This bias never labels itself in a paper; it has to be inferred from how many people received the reference standard, which is why the STARD diagram exists.

Splitting the same dataset by disease severity, with the same test and the same cut-off, leaves the two strata with sensitivities differing by nearly threefold. What does that show?

Show the answer and why

Correct answer: It shows sensitivity depends on which patients were enrolled. In the severe stratum it is 0.815, because sicker people have higher marker concentrations and cross the same cut-off more often

Sensitivity and specificity do not change with prevalence, which is where the 0.197 answer conflates two things. What they do depend on is who is being tested: 0.286 and 0.815 come from one test at one cut-off, and the only difference is that one stratum is mild and the other severe. Nor is this instability in the measurement, since sicker people have higher marker concentrations and cross the cut-off more often, a predictable direction rather than noise. This is exactly why a two-gate design of confirmed cases against healthy volunteers produces beautiful accuracy that cannot be used: it answers whether the test separates severe disease from health, whereas the clinical question is whether it helps with a patient whose presentation is unclear. Note too that the severe stratum contains barely a dozen people without the disease, so its specificity interval is wide enough to carry almost no information.

The sensitivity interval in the mild stratum is so wide it carries almost no information. Which number sets that width?

Show the answer and why

Correct answer: The 14 people with the disease, because the denominator of sensitivity holds only diseased people

Sensitivity is the proportion of diseased people the test picks up, so its denominator contains only diseased people. The mild stratum has 71 people in it, but only 14 of them have the disease, and those 14 set the precision; the remaining 57 affect specificity alone. A study of a thousand people with twenty confirmed cases therefore has the sensitivity precision of twenty people. That is why a diagnostic accuracy study should report the number of cases and the number of controls rather than the total, since a large total may have nothing to do with the measure you care about.

In the demonstration STARD diagram 822 people had the index test, 455 had the reference standard, and the 2 by 2 table is built on 441. What is the gap between 455 and 441?

Show the answer and why

Correct answer: The 14 people with an inconclusive index test - they did have the reference standard, but inconclusive is neither a positive nor a negative reading

822 minus 455 is 367, the people who never had the reference standard and who left one layer earlier. 455 minus 441 is 14, which is the gap this question asks about: people who did have the reference standard but have no cell to sit in. Both groups end up outside the 2 by 2 table for entirely different reasons, and since the sensitivity and specificity a paper reports have a denominator of 441, the reader has to do the subtraction to see who is missing. The 354 is the unverified box on the negative branch, part of the 367 and the place where verification bias lives. Folding inconclusive results into either the positive or the negative column would let a group that could not be classified at all pull that column's accuracy around.

In the demonstration diagram nearly all 214 test-positive people had the reference standard, while only some of the 590 test-negative ones did. What does that asymmetry mean?

Show the answer and why

Correct answer: It means only 236 people on the negative branch were actually verified, and the rest were assumed disease free rather than confirmed to be

On the positive branch 205 of 214 had the reference standard; on the negative branch only 236 of 590 did. That asymmetry is partial verification: the unverified negatives are written into the table as disease free, and some of them are false negatives. The second statement assumes the positive branch already captures everyone with the disease, but a test that caught everyone would not need an accuracy study. The third compares raw counts: the positive branch misses a single-digit number while the negative branch misses more than three hundred, and the degree of asymmetry is exactly the point. Reading a paper calls for the same arithmetic, subtracting the number who received the reference standard from the number enrolled and then asking how the remainder were handled.

The sensitivity in the main analysis comes with a 95% confidence interval. What does its width tell the reader?

Show the answer and why

Correct answer: The lower bound for sensitivity is only 0.481, so these data cannot rule out missing one of every two diseased people

0.634 is the point estimate and it says nothing about what these data can exclude. The lower bound of 0.481 is the number to read: this dataset is compatible with fewer than half of diseased people being detected. The 0.700 is the lower bound for specificity, whose denominator is the people without the disease, and since this cohort has more of those than of cases, the specificity interval is narrower to begin with, so the two measures are not equally precise. Reporting a sensitivity without its interval presents an estimate resting on a few dozen people as though it were a settled number.

Watch next

Cohort Studies: A Brief Overview
ENTerry Shaneyfelt· 6 minA diagnostic accuracy study is a cross-sectional cohort design at heart. Settle the design vocabulary first.

Sources and licences

Report a content problem

The statistics on this site are written by AI and reviewed by AI; a human only spot-checks. What you can see may be what we cannot.

The more specific, the more fixable — e.g. which sentence disagrees with which textbook or paper.

Needed only if you want a reply; reports without it are still read.

Sent along with your report

These are attached automatically. You can drop any of them.