AdvancedIndependently reviewed, not yet spot-checked by a human

Bland-Altman analysis and the limits of agreement

Saying two measurement methods are highly correlated conveys almost nothing. This page uses Bland and Altman's own blood pressure data: first the correlation coefficient talks you into it, then the same 85 people's differences show how wide the limits really are. Then how to handle replicate measurements (what collapses is the precision, not the limits), how to check for proportional bias, and how to translate a limit in mmHg into a clinical decision.

What this page is here to answer

Results sections often read like this: the new method correlated strongly with the reference method, demonstrating good agreement. The first half is a statistic; the second half is a wish, and no inference connects them. The data on this page will make that correlation coefficient look entirely defensible, and then show you the same people’s differences.

Agreement asks something very concrete: if I switch this patient from method A to method B, how different will the number be, and is that difference big enough to change what I do. A correlation coefficient cannot answer it, because it measures how similarly the two methods rank patients, not how far apart two readings on one patient are.

After this page, four questions should be available to you whenever you meet a method-comparison paper:

  1. How large is the bias, and in which direction—does the new method read higher or lower than the reference, and by how much.
  2. How wide are the limits of agreement—converted into clinical units: how many clinically material differences wide is that.
  3. How uncertain are the limits themselves—the limits are estimates with their own confidence intervals, and mishandled replicates usually break this, not the limits.
  4. Do the differences change with the size of the measurement—if they do, one fixed pair of limits cannot serve the whole range.

The data: Bland and Altman’s own blood pressure set

This page uses MethComp::sbp: 765 readings from 85 subjects, each measured 3 times by each of 3 methods, in mmHg. It is the teaching dataset Bland and Altman used themselves, and the three methods are:

CodeWhat it isIts role on this site
Jhuman observer J, sphygmomanometerthe reference method on this page
Rhuman observer R, sphygmomanometerreserved for the inter-observer reliability page
Ssemi-automatic machinethe test method on this page

This page compares human observer J against semi-automatic machine S, and the difference is always defined as S − J: a positive value means the machine read higher than the human. That gives 255 replicate-level pairs, or 85 pairs of per-subject means.

The seduction: the scatter plot first, then the same people’s differences

Two panels side by side, both using the same 85 subjects and each subject's mean of three readings. The left panel is titled A. Correlation -- looks like agreement; the x axis is observer J's mean systolic pressure in mmHg and the y axis is machine S's mean systolic pressure in mmHg, on identical scales. The points form a clearly rising cloud and a dashed line marks the line of equality. A legend at top left prints r = 0.82, a 95% confidence interval of 0.73 to 0.88, p < 0.0001, and n = 85 subjects; a legend at bottom right labels the dashed line as the line of equality. Most points sit above the line of equality, meaning the machine reads higher, but on this scale that is not obvious. The right panel is titled B. Bland-Altman -- the same 85 subjects; the x axis becomes the mean of the two methods in mmHg and the y axis becomes machine S minus observer J in mmHg. A thick solid line marks the bias at 15.6 and is labelled bias; two dashed lines at 56.7 and -25.4 are labelled +1.96 SD and -1.96 SD, and a thin grey line marks zero. The points are spread widely between the dashed lines, tens of mmHg from top to bottom, with the largest single subject difference reaching 105 mmHg; a line of text at bottom left reads limits span 82 mmHg.
Left: the scatter plot and correlation for 85 subjects, which looks like agreement. Right: exactly the same data plotted as difference against mean, where the limits of agreement span 82.1 mmHg. Not one number differs between the two panels; only the question does.Plotting script figures/scripts/B9-03-bland-altman.R

The left panel gives r = 0.817 (95% CI 0.732 to 0.878, p < 0.0001), so r² = 0.668. Computed on all 255 replicate-level pairs instead, it is 0.796.

No clinical paper would refuse to call that a strong correlation—the p value is not remotely in question and the lower confidence bound still sits above 0.73. And it contributes nothing to the question of whether this machine can replace manual measurement.

Why the x axis is the mean of the two methods

The x axis of the right panel is the mean of the two methods, not the reference method. This is not a matter of taste: the difference is negatively correlated with the reference method by construction (whenever the reference reading happens to be high, the difference is pulled down), so plotting difference against reference manufactures a downward slope and makes you see proportional bias that is not there. The mean of the two is the best available approximation to the truth, and it flattens that artefact out.

The three lines on the difference plot come from: the mean difference, that is a bias of +15.62 mmHg (95% CI +11.54 to +19.70); and the bias plus and minus 1.96 times the SD of the differences, 20.95 mmHg, giving -25.4 and +56.7. That is a span of 82.1 mmHg.

Which means: swap this machine in, and an individual patient’s reading may come out 25 mmHg below the manual measurement or 57 mmHg above it—and that is still only the range holding 95% of people.

How wide is too wide: three ways to translate the limits

The limits are a statistic; whether they are good enough is not a statistical question. Three routes turn them into something judgeable, and this dataset supports all three.

First, convert to a clinically material difference. For systolic pressure, 10 mmHg is the usual order of magnitude that changes management. The span of the limits equals 8.2 such differences. The raw distribution of the replicate pairs says the same thing: 62.7% of pairs differ by more than 10 mmHg, 23.5% by more than twice that, and the largest single pair differs by 109 mmHg.

Second, ask whether it flips decisions. Take 140 mmHg as the line. Observer J puts 24 people above it, machine S puts 36. Compared person by person, 14 subjects (16.5%) cross that line purely because of which method measured them—13 of them because the machine called them high and the human did not, 1 in the other direction. How a cut-point is chosen, and how to read misclassification near it, is on choosing a cut-point.

Third, find an external regulatory anchor. ISO 81060-2 criterion 1 requires a non-invasive blood pressure device to differ from the reference by no more than 5 mmHg on average, with an SD of differences no greater than 8 mmHg. This dataset observes a mean difference of 15.62 and an SD of 20.95 mmHg, failing both criteria by a wide margin.

The control condition: same people, same occasion, two human observers

The dataset carries a third method: a second human observer, R. Pair J with R and nothing changes except who the other party is—same sample, same occasion, same arithmetic:

The pairBias (mmHg)SD of differences (mmHg)Limits (mmHg)Span (mmHg)
Machine S against observer J (this page)+15.6220.95-25.4 to +56.782.1
Observer J against observer R+0.092.26-4.3 to +4.58.9

Absolute-agreement ICC between the two human observers is 0.997, and their limits span 8.9 mmHg—roughly 11% of the human-versus-machine span, or put the other way, the machine pair is about 9 times wider.

That control rules out a whole class of excuses. The limits are not wide because blood pressure fluctuates, nor because subjects were uncooperative, nor because 85 people is too few, nor because Bland-Altman is harsh on everyone. Same people, same afternoon, same formulas: two humans reading the same sphygmomanometer differ with an SD of only 2.26 mmHg. What is wide is this machine.

Replicate measurements: the precision collapses, the limits do not

Two difference-against-mean panels side by side sharing one pair of axes: x is the mean of the two methods in mmHg and y is machine S minus observer J in mmHg. The left panel is titled A. All 255 pairs treated as independent, with two lines of small text beneath the title giving SD 20.4, limits -24.3 to 55.5, span 80 mmHg, and noting that the shading is the 95% confidence interval of each limit, 4.3 mmHg either side, with n taken as 255. The panel plots all 255 replicate pairs, a thick solid bias line, two dashed limit lines, and a pale horizontal shaded band around each limit. The right panel is titled B. Bland and Altman (1999) for replicates, with small text giving SD 20.9, limits -25.4 to 56.7, span 82 mmHg, and shading of 6.6 mmHg either side with n = 85 subjects; it plots the 85 per-subject means. The dashed limit lines in the two panels sit at almost the same height, hard to tell apart by eye; the difference is in the shaded bands, which are visibly thicker on the right. The right panel also carries a third pair of limits as dotted lines, computed from per-subject means with no correction, sitting clearly inside the dashed ones, and a legend at top left says that set is 10% too narrow.
Three sets of limits from one dataset. The dashed lines in the two panels are almost in the same place (the spans differ by 2.8%); what differs is the thickness of the shaded bands—treating 85 subjects as 255 independent observations narrows the confidence interval of each limit by 35%. The dotted lines on the right are the uncorrected per-subject-mean version, and that is the set that really is too narrow.Plotting script figures/scripts/B9-03-bland-altman.R

Every subject was measured 3 times, so the question of how many observations there are has three answers, and the three answers give three different sets of limits:

ApproachnSD of differencesLimitsSpan95% CI half-width of each limit
A. all 255 pairs treated as independent25520.37-24.3 to +55.579.84.30
B. per-subject mean differences, no correction8518.93-21.5 to +52.774.27.01
C. Bland and Altman (1999) replicate correction8520.95-25.4 to +56.782.16.60

All values in mmHg. C is what this page uses. Put the three rows side by side and the conclusion runs against most people’s intuition:

Treating all 255 pairs as independent barely moves the limits at all. A span of 79.84 against the correct 82.12 is 2.77% narrow, or 2.28 mmHg in absolute terms. Drawn on a plot, the two pairs of dashed lines are not distinguishable by eye.

What collapses is the precision of the limits. The 95% confidence interval of each limit shrinks from the correct 6.60 mmHg either side to 4.30 mmHg either side, 34.9% narrower. The reason is direct: the effective sample size is inflated 3-fold—there are only 85 independent people in the data, but the standard errors are computed as though there were 255 independent observations.

The approach that really is badly too narrow is the third one. Take each subject’s mean of 3 readings, difference those, and apply no correction: the span is 74.22 mmHg, 9.62% narrow, 7.90 mmHg missing. This is what section 5.2 of Bland and Altman (1999) was warning about: averaging shrinks each method’s own measurement error, so the limits you report describe a mean of three against a mean of three, while a patient in clinic is measured once.

Why the limits did not collapse too

Because J and S read the same subject at the same time. Whatever the true pressure was at that instant is shared by both methods, so it subtracts out within each occasion’s difference. The variance of the differences therefore contains almost nothing of where this person’s pressure happened to be, and repeating three times does not conjure any of it back—so treating the replicates as independent estimates an SD close to the right answer.

The correct procedure splits the variance of the difference and rebuilds it. With mm replicates per subject, dˉ\bar d the per-subject mean difference, and sw2s_w^2 each method’s own within-subject variance:

sd2  =  var(dˉ)  +  (11m)sw,ref2  +  (11m)sw,test2s_d^2 \;=\; \operatorname{var}(\bar d) \;+\; \left(1 - \frac{1}{m}\right) s_{w,\text{ref}}^2 \;+\; \left(1 - \frac{1}{m}\right) s_{w,\text{test}}^2

For this dataset: mm is 3, so the correction factor is 0.667; the SD of the per-subject mean differences is 18.93 mmHg; observer J’s own within-subject SD across three readings is 6.12 mmHg and machine S’s is 9.12 mmHg (on 170 degrees of freedom each). Rebuilt, that is 20.95 mmHg, which is row C above.

Notice the relationship between those two within-subject SDs in passing: the machine repeating itself three times varies more than the human does (9.12 against 6.12 mmHg). A device sold on removing observer error is not, on its own repeatability, steadier than the observer.

Proportional bias: checked, and none detected

Two scatter panels side by side, each with a fitted regression line and its confidence band. The left panel is titled A. Difference against magnitude; x is the mean of the two methods in mmHg and y is machine S minus observer J in mmHg, and a line of small text under the title gives a slope of +0.036 mmHg per mmHg with a 95% confidence interval of −0.103 to +0.174, p = 0.610. The panel shows 85 points, a grey zero line, and a thick fitted line with a pale confidence band; the fitted line is almost horizontal. A line of text at top left reads: the fitted line is flat, no proportional bias detected. The right panel is titled B. Same check on the log scale; x is the mean of the log measurements in log mmHg and y is log of machine S minus log of observer J, with small text giving a slope of −0.064, 95% confidence interval −0.196 to +0.068, p = 0.338. It likewise shows points, a zero line, an almost horizontal fitted line and its band, plus two dashed lines marking the limits of agreement on the log scale, labelled +1.96 SD equals a ratio of 1.49 and -1.96 SD equals a ratio of 0.85. A line of text at top left says the machine reads 12.6% higher on average.
Regressing the difference on the mean asks whether the differences change with the size of the measurement. Left: the original scale, slope +0.036 mmHg per mmHg, p = 0.610. Right: the same check on the log scale, where the limits come back as ratios. Neither detects proportional bias.Plotting script figures/scripts/B9-03-bland-altman.R

A single pair of limits carries an implicit premise: the size of the difference is unrelated to the size of the measurement. If people with higher pressures differ more, that premise fails, and one pair of limits stretched across the whole range is too wide at the low end and too narrow at the high end.

The check is a linear regression of the difference on the mean, read off the slope. On per-subject means the slope here is +0.0356 mmHg/mmHg (95% CI −0.1028 to +0.1739, p = 0.610); on all 255 replicate pairs it is +0.0514 (p = 0.226). The point estimate says that a 10 mmHg rise in the measurement shifts the difference by +0.36 mmHg—negligible next to 10 mmHg.

So the finding on this page is: no proportional bias was detected in these data, and the limits are reported as one fixed pair.

The right panel demonstrates the other common move: take logs of both methods and repeat. That is the standard handling when differences scale with the magnitude, and its advantage is that the limits come back as ratios rather than absolute amounts. Here the log-scale slope is −0.0640 (p = 0.338), again with no trend detected; translated back, the machine reads 12.6% higher on average, with limits at ratios of 0.851 to 1.491, that is -14.9% to +49.1%.

Running it yourself

data(sbp, package = "MethComp")

# The data are long: meth (J/R/S), item (subject), repl (which reading), y (mmHg).
# Assert one row per (meth, item, repl) first — otherwise the reshape below
# silently keeps the first match instead of erroring. This line is load-bearing.
stopifnot(!anyDuplicated(sbp[c("meth", "item", "repl")]))

w <- reshape(sbp[sbp$meth %in% c("J", "S"), ],
             direction = "wide", idvar = c("item", "repl"),
             timevar = "meth")
w$diff <- w$y.S - w$y.J          # difference is machine minus human
w$avg  <- (w$y.S + w$y.J) / 2

# Per-subject means (85 subjects)
m <- aggregate(cbind(diff, avg, y.J, y.S) ~ item, data = w, FUN = mean)

# 1. The seduction: the correlation. High, and unrelated to agreement.
cor.test(m$y.J, m$y.S)

# 2. Limits of agreement with the Bland & Altman (1999) replicate correction.
#    Taking sd(w$diff) over the 255 pairs treats 85 people as 255 independent
#    observations: the limits come out only a few percent narrow, but the
#    confidence interval of each limit comes out a third too narrow.
mm <- 3
within_var <- function(v) {
  # each method's own within-subject variance, pooled by degrees of freedom
  a <- aggregate(v ~ item, data = w, FUN = function(x) c(var(x), length(x) - 1))
  sum(a$v[, 1] * a$v[, 2]) / sum(a$v[, 2])
}
n      <- nrow(m)
df_w   <- n * (mm - 1)
k      <- 1 - 1 / mm
vJ     <- within_var(w$y.J)
vS     <- within_var(w$y.S)
v_dbar <- var(m$diff)
s2     <- v_dbar + k * (vJ + vS)
bias   <- mean(m$diff)
z      <- qnorm(0.975)
loa    <- bias + c(-1, 1) * z * sqrt(s2)

# 3. Confidence intervals for the limits. s2 is a sum of three independent
#    components, so its variance is the sum of three Var(s2) = 2 s2^2 / df
#    terms, and the limit's variance is then
#    Var(LoA) = Var(bias) + z^2 * Var(s), with Var(s) = Var(s2) / (4 s2).
#    The textbook sqrt(s2 * (1/n + z^2 / (2(n-1)))) assumes every observation
#    is independent, which is exactly what is false here: it comes out about a
#    millimetre of mercury wider than the table above. The denominator is
#    subjects, not pairs.
var_s2 <- 2 * v_dbar^2 / (n - 1) + k^2 * 2 * vJ^2 / df_w + k^2 * 2 * vS^2 / df_w
se     <- sqrt(v_dbar / n + z^2 * var_s2 / (4 * s2))
tq     <- qt(0.975, n - 1)
rbind(lower = loa[1] + c(-1, 1) * tq * se,
      upper = loa[2] + c(-1, 1) * tq * se)

# 4. Proportional bias: regress difference on mean. The x variable must be the
#    mean of the two — plotting against the reference manufactures a slope.
summary(lm(diff ~ avg, data = m))

# 5. The log-scale version: the limits come back as ratios, not mmHg. Fit it on
#    the per-subject means, as in step 4 — run over the 255 single pairs it
#    returns a different slope, which is not the one the table reports.
lm(log(y.S) - log(y.J) ~ I((log(y.S) + log(y.J)) / 2), data = m)

# Exported for the Python block below
write.csv(sbp, "sbp.csv", row.names = FALSE)

Verified against R 4.6.0 with MethComp 1.30.2. MethComp is used only to obtain the data; every calculation is base R, because the point of this page is that the formulas stay visible.

Reading the report

A method-comparison paper usually gives you one Bland-Altman plot and a paragraph. Six places are worth checking one by one:

  1. Is the direction of the difference stated. A minus B and B minus A have opposite signs for the bias. The y axis label should say who is subtracted from whom; a plot labelled only difference cannot be quoted.
  2. Do the limits come with confidence intervals. The limits are estimates. The upper limit here is +56.7 and its 95% CI reaches +63.3—if the authors want to argue that the limits fall within an acceptable range, the outer edge of that interval is what matters, not the point estimate.
  3. How were the replicates handled. Replicates present and n reported as the number of pairs (here that would be 255 rather than 85) means the confidence intervals of the limits are too narrow; per-subject means with no correction means the limits themselves are too narrow. A Methods section that says nothing leaves both possible.
  4. Was proportional bias checked, and how. There should be a regression of difference on the mean of the two methods. If it was regressed on the reference method, that slope cannot be trusted.
  5. Is there a pre-specified acceptable difference. Without that number, acceptable limits is just the authors’ impression. The best version cites an external standard, as with ISO 81060-2 criterion 1 here.
  6. Does the sample cover enough of the range. Limits extrapolate only as far as the data go. A cohort clustered in the normal range says nothing about how the machine behaves above 140 mmHg.

Common misuses

MisuseWhy it is wrong
Arguing agreement from r or R²It measures rank association and rises with how spread out the subjects are; a uniform offset of 16 mmHg leaves it untouched
Using a paired t test to prove the two methods do not differIt tests whether the mean difference is zero and says nothing about the spread of individual differences; non-significance more often means too small a sample
Plotting the difference against the reference methodThe reference method’s error enters both axes and manufactures a downward slope
Replicates present, all pairs treated as independentThe limits here are only 2.8% narrow, but the confidence interval of each limit is 35% narrow — and the claim of acceptable limits rests on that interval
Per-subject mean differences with no correctionAveraging shrinks each method’s own error; here it drops 7.9 mmHg of span, while a patient in clinic is measured once
Reporting limits with no pre-specified acceptable differenceThe limits are a statistic; good enough needs an external standard or a clinical threshold
Turning a non-significant slope into “the difference is unrelated to the magnitude”Not detected is not absent; the interval here still allows a tilt of 10 to 17 mmHg across the range
Extrapolating the limits beyond the measured rangeThe limits describe the range that was measured; another population needs re-estimation
Wide limits, concluding the methods are interchangeableThe span here is 82 mmHg, 8.2 clinically material differences, and 16.5% of people cross 140 mmHg
Offering an ICC as evidence of agreementAn ICC is a ratio whose denominator contains the between-subject variance; the same device looks excellent in a widely spread population

How this page relates to the others

  • The reliability side of the same data—see the intraclass correlation coefficient. That page uses the J against R pair this one deliberately avoids, and together they show that agreement and reliability are two different questions asked of one set of numbers.
  • Agreement for categorical outcomes—see kappa and weighted kappa. When the outcome is a category rather than a quantity, a difference has no meaning and the tooling changes.
  • The downstream consequences of measurement error—see measurement error. The spread of differences measured here is what that page discusses, seen before it enters a regression.
  • That regression of difference on mean—see linear regression. The proportional-bias check is just that model applied to a particular pair of variables.
  • Dependence between repeated measurements—see clustered data. The correction formula here is a special case of that page’s variance decomposition.
  • Cut-points and flipped decisions—see choosing a cut-point. The second translation route above uses that page’s framing.
  • Clinical context—see the design chapter diagnostic accuracy studies.

Regenerating every number on this page

/opt/homebrew/bin/Rscript figures/scripts/B9-03-bland-altman.R

Read the figure

The answer comes from the same statistical output that produced this page's figures, not from a number typed in beside them.

Machine S and observer J measure systolic blood pressure in the same people, and the correlation is 0.817 (95% CI 0.732 to 0.878). Does that support replacing the observer with the machine?

Show the answer and why

Correct answer: No. Plot the differences for the same people and the limits of agreement span 82.1 mmHg

A correlation measures whether two sets of numbers rise and fall together, not whether they produce the same number: add a constant to every machine reading and the correlation does not move at all, while agreement is destroyed. Plot the differences for the same people and the limits span 82.1 mmHg, so swapping in this machine can put an individual patient twenty-odd mmHg below the observer or fifty-odd above. 15.6 mmHg is the mean difference, which gives the systematic offset and says nothing about how far one patient can be out. 6.1 mmHg is the within-subject standard deviation of the observer repeating himself, which is about how one method reproduces itself. And a unitless correlation cannot be compared in size with a quantity in mmHg in the first place.

On the difference plot the middle line sits at the mean difference of 15.62 mmHg, and the standard deviation of the differences is 20.95 mmHg. Where is the upper line, and what does it mean?

Show the answer and why

Correct answer: At 56.68, the upper end of the band between the two lines that holds 95% of individual differences

The upper line is the mean difference plus roughly two standard deviations of the differences, which puts it at 56.68, and 95% of individual differences fall between the two lines. 19.70 is the upper end of the confidence interval for the mean difference, which says how precisely the average offset is known and shrinks towards nothing as the sample grows; reporting it as a limit claims the machine and the observer barely differ. 63.28 is the upper end of the confidence interval for the limit itself, which answers how uncertain the position of that line is and belongs alongside the limit, but the limit is still 56.68. All three print on the same figure, and mixing them answers a question about individual patients with an interval about an average.

Each subject was measured three times. Treating all 255 pairs as independent gives limits spanning 79.84 mmHg, while the method built for replicates gives 82.12 mmHg. What does treating replicates as independent actually cost?

Show the answer and why

Correct answer: The 95% confidence interval for each limit comes out as 4.30 mmHg either side, narrower than it should be, because the effective sample size is inflated

The limits themselves barely move, 79.84 against 82.12, because the two methods read the same subject at the same moment, so whatever the blood pressure happened to be cancels out of the difference. What collapses is the precision of the limits: the confidence interval goes from a correct 6.60 mmHg either side to 4.30 mmHg, because the data hold eighty-five independent people while the standard error was computed as though there were 255 independent observations. 74.22 mmHg is the genuinely too-narrow set, and it comes from a different approach, differencing each subject's mean of three readings with no correction, which describes a mean of three against a mean of three while every patient in clinic is measured once. A narrower interval claims more certainty than the data support, which pushes a device that failed towards passing.

Regressing the difference on the mean of the two methods gives a slope of 0.0356 mmHg/mmHg per subject, 95% CI -0.1028 to 0.1739, p = 0.610. What may be written?

Show the answer and why

Correct answer: The upper limit of 0.1739 multiplied by the hundred-plus mmHg range of these data is still not small, so all that may be written is that no trend was detected

Failing to reach significance usually means the sample was too small, not that the effect is zero. The upper limit of 0.1739 mmHg/mmHg across the hundred-plus mmHg range covered here leaves a tilt of a dozen or so mmHg between the ends of the range compatible with the data, and that is not a range small enough to declare absent. The correct sentence is that no trend of the difference with magnitude was detected, not that the difference is independent of magnitude; the second claim needs an equivalence margin fixed in advance. 0.6104 is the p value, and not detected is not the same as not there. 0.0514 is the slope on all the single-replicate pairs, and two flat estimates pointing the same way corroborate each other rather than failing.

Taking 140 mmHg as the cut-off, 24 of the 85 people are high by observer J and 36 are high by machine S. What does that say about replacing the observer with the machine?

Show the answer and why

Correct answer: Compared person by person, 14 cross the line purely because the measurement method changed

Comparing totals hides it. Person by person, 14 cross 140 mmHg purely because the method changed, 13 of them high on the machine and not on the observer and one the other way. 13 is that one-directional count, not the difference between the two totals. 24 is the number the observer called high, and whether the people the machine adds are truly hypertensive is a question these data cannot answer, because there is no third-party reference standard here; all they show is who changes side. Translating limits into clinical terms means asking whether decisions flip, not only reading the mean difference.

Same people, same afternoon, with only the identity of the other party swapped: two human observers against each other give a standard deviation of differences of just 2.26 mmHg on single readings. What does that control rule out?

Show the answer and why

Correct answer: Their limits span 8.86 mmHg, so the wide limits against the machine cannot be blamed on blood pressure being variable, on too few people, or on the method being harsh

Same people, same afternoon, same formulas, with only the identity of the other party changed. The two human observers span 8.86 mmHg on single readings, so the width against the machine is not blood pressure moving, not uncooperative subjects, and not a method that is harsh on everyone: the machine is what is wide. 30.63 mmHg is the standard deviation of the subject means, which describes how much the people differ from one another rather than any set of limits. 5.24 mmHg comes from differencing each person's mean of three readings, and averaging shrinks each method's own measurement error, so that figure describes a mean against a mean while every patient in clinic is measured once.

Sources and licences

This page is original writing

Report a content problem

The statistics on this site are written by AI and reviewed by AI; a human only spot-checks. What you can see may be what we cannot.

The more specific, the more fixable — e.g. which sentence disagrees with which textbook or paper.

Needed only if you want a reply; reports without it are still read.

Sent along with your report

These are attached automatically. You can drop any of them.