Bland-Altman analysis and the limits of agreement
Saying two measurement methods are highly correlated conveys almost nothing. This page uses Bland and Altman's own blood pressure data: first the correlation coefficient talks you into it, then the same 85 people's differences show how wide the limits really are. Then how to handle replicate measurements (what collapses is the precision, not the limits), how to check for proportional bias, and how to translate a limit in mmHg into a clinical decision.
What this page is here to answer
Results sections often read like this: the new method correlated strongly with the reference method, demonstrating good agreement. The first half is a statistic; the second half is a wish, and no inference connects them. The data on this page will make that correlation coefficient look entirely defensible, and then show you the same people’s differences.
Agreement asks something very concrete: if I switch this patient from method A to method B, how different will the number be, and is that difference big enough to change what I do. A correlation coefficient cannot answer it, because it measures how similarly the two methods rank patients, not how far apart two readings on one patient are.
After this page, four questions should be available to you whenever you meet a method-comparison paper:
- How large is the bias, and in which direction—does the new method read higher or lower than the reference, and by how much.
- How wide are the limits of agreement—converted into clinical units: how many clinically material differences wide is that.
- How uncertain are the limits themselves—the limits are estimates with their own confidence intervals, and mishandled replicates usually break this, not the limits.
- Do the differences change with the size of the measurement—if they do, one fixed pair of limits cannot serve the whole range.
The data: Bland and Altman’s own blood pressure set
This page uses MethComp::sbp: 765 readings from 85 subjects, each measured 3 times by each of 3 methods, in mmHg. It is the teaching dataset Bland and Altman used themselves, and the three methods are:
| Code | What it is | Its role on this site |
|---|---|---|
J | human observer J, sphygmomanometer | the reference method on this page |
R | human observer R, sphygmomanometer | reserved for the inter-observer reliability page |
S | semi-automatic machine | the test method on this page |
This page compares human observer J against semi-automatic machine S, and the difference is always defined as S − J: a positive value means the machine read higher than the human. That gives 255 replicate-level pairs, or 85 pairs of per-subject means.
The seduction: the scatter plot first, then the same people’s differences
figures/scripts/B9-03-bland-altman.RThe left panel gives r = 0.817 (95% CI 0.732 to 0.878, p < 0.0001), so r² = 0.668. Computed on all 255 replicate-level pairs instead, it is 0.796.
No clinical paper would refuse to call that a strong correlation—the p value is not remotely in question and the lower confidence bound still sits above 0.73. And it contributes nothing to the question of whether this machine can replace manual measurement.
Why the x axis is the mean of the two methods
The x axis of the right panel is the mean of the two methods, not the reference method. This is not a matter of taste: the difference is negatively correlated with the reference method by construction (whenever the reference reading happens to be high, the difference is pulled down), so plotting difference against reference manufactures a downward slope and makes you see proportional bias that is not there. The mean of the two is the best available approximation to the truth, and it flattens that artefact out.
The three lines on the difference plot come from: the mean difference, that is a bias of +15.62 mmHg (95% CI +11.54 to +19.70); and the bias plus and minus 1.96 times the SD of the differences, 20.95 mmHg, giving -25.4 and +56.7. That is a span of 82.1 mmHg.
Which means: swap this machine in, and an individual patient’s reading may come out 25 mmHg below the manual measurement or 57 mmHg above it—and that is still only the range holding 95% of people.
How wide is too wide: three ways to translate the limits
The limits are a statistic; whether they are good enough is not a statistical question. Three routes turn them into something judgeable, and this dataset supports all three.
First, convert to a clinically material difference. For systolic pressure, 10 mmHg is the usual order of magnitude that changes management. The span of the limits equals 8.2 such differences. The raw distribution of the replicate pairs says the same thing: 62.7% of pairs differ by more than 10 mmHg, 23.5% by more than twice that, and the largest single pair differs by 109 mmHg.
Second, ask whether it flips decisions. Take 140 mmHg as the line. Observer J puts 24 people above it, machine S puts 36. Compared person by person, 14 subjects (16.5%) cross that line purely because of which method measured them—13 of them because the machine called them high and the human did not, 1 in the other direction. How a cut-point is chosen, and how to read misclassification near it, is on choosing a cut-point.
Third, find an external regulatory anchor. ISO 81060-2 criterion 1 requires a non-invasive blood pressure device to differ from the reference by no more than 5 mmHg on average, with an SD of differences no greater than 8 mmHg. This dataset observes a mean difference of 15.62 and an SD of 20.95 mmHg, failing both criteria by a wide margin.
The control condition: same people, same occasion, two human observers
The dataset carries a third method: a second human observer, R. Pair J with R and nothing changes except who the other party is—same sample, same occasion, same arithmetic:
| The pair | Bias (mmHg) | SD of differences (mmHg) | Limits (mmHg) | Span (mmHg) |
|---|---|---|---|---|
| Machine S against observer J (this page) | +15.62 | 20.95 | -25.4 to +56.7 | 82.1 |
| Observer J against observer R | +0.09 | 2.26 | -4.3 to +4.5 | 8.9 |
Absolute-agreement ICC between the two human observers is 0.997, and their limits span 8.9 mmHg—roughly 11% of the human-versus-machine span, or put the other way, the machine pair is about 9 times wider.
That control rules out a whole class of excuses. The limits are not wide because blood pressure fluctuates, nor because subjects were uncooperative, nor because 85 people is too few, nor because Bland-Altman is harsh on everyone. Same people, same afternoon, same formulas: two humans reading the same sphygmomanometer differ with an SD of only 2.26 mmHg. What is wide is this machine.
Replicate measurements: the precision collapses, the limits do not
figures/scripts/B9-03-bland-altman.REvery subject was measured 3 times, so the question of how many observations there are has three answers, and the three answers give three different sets of limits:
| Approach | n | SD of differences | Limits | Span | 95% CI half-width of each limit |
|---|---|---|---|---|---|
| A. all 255 pairs treated as independent | 255 | 20.37 | -24.3 to +55.5 | 79.8 | 4.30 |
| B. per-subject mean differences, no correction | 85 | 18.93 | -21.5 to +52.7 | 74.2 | 7.01 |
| C. Bland and Altman (1999) replicate correction | 85 | 20.95 | -25.4 to +56.7 | 82.1 | 6.60 |
All values in mmHg. C is what this page uses. Put the three rows side by side and the conclusion runs against most people’s intuition:
Treating all 255 pairs as independent barely moves the limits at all. A span of 79.84 against the correct 82.12 is 2.77% narrow, or 2.28 mmHg in absolute terms. Drawn on a plot, the two pairs of dashed lines are not distinguishable by eye.
What collapses is the precision of the limits. The 95% confidence interval of each limit shrinks from the correct 6.60 mmHg either side to 4.30 mmHg either side, 34.9% narrower. The reason is direct: the effective sample size is inflated 3-fold—there are only 85 independent people in the data, but the standard errors are computed as though there were 255 independent observations.
The approach that really is badly too narrow is the third one. Take each subject’s mean of 3 readings, difference those, and apply no correction: the span is 74.22 mmHg, 9.62% narrow, 7.90 mmHg missing. This is what section 5.2 of Bland and Altman (1999) was warning about: averaging shrinks each method’s own measurement error, so the limits you report describe a mean of three against a mean of three, while a patient in clinic is measured once.
Why the limits did not collapse too
Because J and S read the same subject at the same time. Whatever the true pressure was at that instant is shared by both methods, so it subtracts out within each occasion’s difference. The variance of the differences therefore contains almost nothing of where this person’s pressure happened to be, and repeating three times does not conjure any of it back—so treating the replicates as independent estimates an SD close to the right answer.
The correct procedure splits the variance of the difference and rebuilds it. With replicates per subject, the per-subject mean difference, and each method’s own within-subject variance:
For this dataset: is 3, so the correction factor is 0.667; the SD of the per-subject mean differences is 18.93 mmHg; observer J’s own within-subject SD across three readings is 6.12 mmHg and machine S’s is 9.12 mmHg (on 170 degrees of freedom each). Rebuilt, that is 20.95 mmHg, which is row C above.
Notice the relationship between those two within-subject SDs in passing: the machine repeating itself three times varies more than the human does (9.12 against 6.12 mmHg). A device sold on removing observer error is not, on its own repeatability, steadier than the observer.
Proportional bias: checked, and none detected
figures/scripts/B9-03-bland-altman.RA single pair of limits carries an implicit premise: the size of the difference is unrelated to the size of the measurement. If people with higher pressures differ more, that premise fails, and one pair of limits stretched across the whole range is too wide at the low end and too narrow at the high end.
The check is a linear regression of the difference on the mean, read off the slope. On per-subject means the slope here is +0.0356 mmHg/mmHg (95% CI −0.1028 to +0.1739, p = 0.610); on all 255 replicate pairs it is +0.0514 (p = 0.226). The point estimate says that a 10 mmHg rise in the measurement shifts the difference by +0.36 mmHg—negligible next to 10 mmHg.
So the finding on this page is: no proportional bias was detected in these data, and the limits are reported as one fixed pair.
The right panel demonstrates the other common move: take logs of both methods and repeat. That is the standard handling when differences scale with the magnitude, and its advantage is that the limits come back as ratios rather than absolute amounts. Here the log-scale slope is −0.0640 (p = 0.338), again with no trend detected; translated back, the machine reads 12.6% higher on average, with limits at ratios of 0.851 to 1.491, that is -14.9% to +49.1%.
Running it yourself
data(sbp, package = "MethComp")
# The data are long: meth (J/R/S), item (subject), repl (which reading), y (mmHg).
# Assert one row per (meth, item, repl) first — otherwise the reshape below
# silently keeps the first match instead of erroring. This line is load-bearing.
stopifnot(!anyDuplicated(sbp[c("meth", "item", "repl")]))
w <- reshape(sbp[sbp$meth %in% c("J", "S"), ],
direction = "wide", idvar = c("item", "repl"),
timevar = "meth")
w$diff <- w$y.S - w$y.J # difference is machine minus human
w$avg <- (w$y.S + w$y.J) / 2
# Per-subject means (85 subjects)
m <- aggregate(cbind(diff, avg, y.J, y.S) ~ item, data = w, FUN = mean)
# 1. The seduction: the correlation. High, and unrelated to agreement.
cor.test(m$y.J, m$y.S)
# 2. Limits of agreement with the Bland & Altman (1999) replicate correction.
# Taking sd(w$diff) over the 255 pairs treats 85 people as 255 independent
# observations: the limits come out only a few percent narrow, but the
# confidence interval of each limit comes out a third too narrow.
mm <- 3
within_var <- function(v) {
# each method's own within-subject variance, pooled by degrees of freedom
a <- aggregate(v ~ item, data = w, FUN = function(x) c(var(x), length(x) - 1))
sum(a$v[, 1] * a$v[, 2]) / sum(a$v[, 2])
}
n <- nrow(m)
df_w <- n * (mm - 1)
k <- 1 - 1 / mm
vJ <- within_var(w$y.J)
vS <- within_var(w$y.S)
v_dbar <- var(m$diff)
s2 <- v_dbar + k * (vJ + vS)
bias <- mean(m$diff)
z <- qnorm(0.975)
loa <- bias + c(-1, 1) * z * sqrt(s2)
# 3. Confidence intervals for the limits. s2 is a sum of three independent
# components, so its variance is the sum of three Var(s2) = 2 s2^2 / df
# terms, and the limit's variance is then
# Var(LoA) = Var(bias) + z^2 * Var(s), with Var(s) = Var(s2) / (4 s2).
# The textbook sqrt(s2 * (1/n + z^2 / (2(n-1)))) assumes every observation
# is independent, which is exactly what is false here: it comes out about a
# millimetre of mercury wider than the table above. The denominator is
# subjects, not pairs.
var_s2 <- 2 * v_dbar^2 / (n - 1) + k^2 * 2 * vJ^2 / df_w + k^2 * 2 * vS^2 / df_w
se <- sqrt(v_dbar / n + z^2 * var_s2 / (4 * s2))
tq <- qt(0.975, n - 1)
rbind(lower = loa[1] + c(-1, 1) * tq * se,
upper = loa[2] + c(-1, 1) * tq * se)
# 4. Proportional bias: regress difference on mean. The x variable must be the
# mean of the two — plotting against the reference manufactures a slope.
summary(lm(diff ~ avg, data = m))
# 5. The log-scale version: the limits come back as ratios, not mmHg. Fit it on
# the per-subject means, as in step 4 — run over the 255 single pairs it
# returns a different slope, which is not the one the table reports.
lm(log(y.S) - log(y.J) ~ I((log(y.S) + log(y.J)) / 2), data = m)
# Exported for the Python block below
write.csv(sbp, "sbp.csv", row.names = FALSE)Verified against R 4.6.0 with MethComp 1.30.2. MethComp is used only to obtain the data; every calculation is base R, because the point of this page is that the formulas stay visible.
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
from scipy import stats
# MethComp is not in the Rdatasets index, so get_rdataset("sbp", "MethComp")
# returns a 404. The R block above already writes it out.
sbp = pd.read_csv("sbp.csv")
assert not sbp.duplicated(["meth", "item", "repl"]).any()
w = (sbp[sbp.meth.isin(["J", "S"])]
.pivot(index=["item", "repl"], columns="meth", values="y")
.reset_index())
w["diff"] = w["S"] - w["J"]
w["avg"] = (w["S"] + w["J"]) / 2
m = w.groupby("item")[["diff", "avg", "J", "S"]].mean()
print(stats.pearsonr(m["J"], m["S"]))
mm = 3
def within_var(col):
g = w.groupby("item")[col]
return (g.var(ddof=1) * (g.count() - 1)).sum() / (g.count() - 1).sum()
n = len(m)
df_w = n * (mm - 1)
k = 1 - 1 / mm
vJ, vS = within_var("J"), within_var("S")
v_dbar = m["diff"].var(ddof=1)
s2 = v_dbar + k * (vJ + vS)
bias = m["diff"].mean()
z = stats.norm.ppf(0.975)
loa = bias + np.array([-1, 1]) * z * np.sqrt(s2)
var_s2 = 2 * v_dbar**2 / (n - 1) + k**2 * 2 * vJ**2 / df_w + k**2 * 2 * vS**2 / df_w
se = np.sqrt(v_dbar / n + z**2 * var_s2 / (4 * s2))
tq = stats.t.ppf(0.975, n - 1)
print(loa, loa[:, None] + np.array([-1, 1]) * tq * se)
print(smf.ols("diff ~ avg", data=m).fit().summary())
lg = m.assign(ld=np.log(m["S"]) - np.log(m["J"]),
la=(np.log(m["S"]) + np.log(m["J"])) / 2)
print(smf.ols("ld ~ la", data=lg).fit().params)The Python side uses pandas and statsmodels with the same structure. sbp is not in the Rdatasets index, so get_rdataset cannot fetch it; export it from R as a CSV first.
Reading the report
A method-comparison paper usually gives you one Bland-Altman plot and a paragraph. Six places are worth checking one by one:
- Is the direction of the difference stated. A minus B and B minus A have opposite signs for the bias. The y axis label should say who is subtracted from whom; a plot labelled only difference cannot be quoted.
- Do the limits come with confidence intervals. The limits are estimates. The upper limit here is +56.7 and its 95% CI reaches +63.3—if the authors want to argue that the limits fall within an acceptable range, the outer edge of that interval is what matters, not the point estimate.
- How were the replicates handled. Replicates present and n reported as the number of pairs (here that would be 255 rather than 85) means the confidence intervals of the limits are too narrow; per-subject means with no correction means the limits themselves are too narrow. A Methods section that says nothing leaves both possible.
- Was proportional bias checked, and how. There should be a regression of difference on the mean of the two methods. If it was regressed on the reference method, that slope cannot be trusted.
- Is there a pre-specified acceptable difference. Without that number, acceptable limits is just the authors’ impression. The best version cites an external standard, as with ISO 81060-2 criterion 1 here.
- Does the sample cover enough of the range. Limits extrapolate only as far as the data go. A cohort clustered in the normal range says nothing about how the machine behaves above 140 mmHg.
Common misuses
| Misuse | Why it is wrong |
|---|---|
| Arguing agreement from r or R² | It measures rank association and rises with how spread out the subjects are; a uniform offset of 16 mmHg leaves it untouched |
| Using a paired t test to prove the two methods do not differ | It tests whether the mean difference is zero and says nothing about the spread of individual differences; non-significance more often means too small a sample |
| Plotting the difference against the reference method | The reference method’s error enters both axes and manufactures a downward slope |
| Replicates present, all pairs treated as independent | The limits here are only 2.8% narrow, but the confidence interval of each limit is 35% narrow — and the claim of acceptable limits rests on that interval |
| Per-subject mean differences with no correction | Averaging shrinks each method’s own error; here it drops 7.9 mmHg of span, while a patient in clinic is measured once |
| Reporting limits with no pre-specified acceptable difference | The limits are a statistic; good enough needs an external standard or a clinical threshold |
| Turning a non-significant slope into “the difference is unrelated to the magnitude” | Not detected is not absent; the interval here still allows a tilt of 10 to 17 mmHg across the range |
| Extrapolating the limits beyond the measured range | The limits describe the range that was measured; another population needs re-estimation |
| Wide limits, concluding the methods are interchangeable | The span here is 82 mmHg, 8.2 clinically material differences, and 16.5% of people cross 140 mmHg |
| Offering an ICC as evidence of agreement | An ICC is a ratio whose denominator contains the between-subject variance; the same device looks excellent in a widely spread population |
How this page relates to the others
- The reliability side of the same data—see the intraclass correlation coefficient. That page uses the J against R pair this one deliberately avoids, and together they show that agreement and reliability are two different questions asked of one set of numbers.
- Agreement for categorical outcomes—see kappa and weighted kappa. When the outcome is a category rather than a quantity, a difference has no meaning and the tooling changes.
- The downstream consequences of measurement error—see measurement error. The spread of differences measured here is what that page discusses, seen before it enters a regression.
- That regression of difference on mean—see linear regression. The proportional-bias check is just that model applied to a particular pair of variables.
- Dependence between repeated measurements—see clustered data. The correction formula here is a special case of that page’s variance decomposition.
- Cut-points and flipped decisions—see choosing a cut-point. The second translation route above uses that page’s framing.
- Clinical context—see the design chapter diagnostic accuracy studies.
Regenerating every number on this page
/opt/homebrew/bin/Rscript figures/scripts/B9-03-bland-altman.RRead the figure
The answer comes from the same statistical output that produced this page's figures, not from a number typed in beside them.
Machine S and observer J measure systolic blood pressure in the same people, and the correlation is 0.817 (95% CI 0.732 to 0.878). Does that support replacing the observer with the machine?
Show the answer and why
Correct answer: No. Plot the differences for the same people and the limits of agreement span 82.1 mmHg
A correlation measures whether two sets of numbers rise and fall together, not whether they produce the same number: add a constant to every machine reading and the correlation does not move at all, while agreement is destroyed. Plot the differences for the same people and the limits span 82.1 mmHg, so swapping in this machine can put an individual patient twenty-odd mmHg below the observer or fifty-odd above. 15.6 mmHg is the mean difference, which gives the systematic offset and says nothing about how far one patient can be out. 6.1 mmHg is the within-subject standard deviation of the observer repeating himself, which is about how one method reproduces itself. And a unitless correlation cannot be compared in size with a quantity in mmHg in the first place.
On the difference plot the middle line sits at the mean difference of 15.62 mmHg, and the standard deviation of the differences is 20.95 mmHg. Where is the upper line, and what does it mean?
Show the answer and why
Correct answer: At 56.68, the upper end of the band between the two lines that holds 95% of individual differences
The upper line is the mean difference plus roughly two standard deviations of the differences, which puts it at 56.68, and 95% of individual differences fall between the two lines. 19.70 is the upper end of the confidence interval for the mean difference, which says how precisely the average offset is known and shrinks towards nothing as the sample grows; reporting it as a limit claims the machine and the observer barely differ. 63.28 is the upper end of the confidence interval for the limit itself, which answers how uncertain the position of that line is and belongs alongside the limit, but the limit is still 56.68. All three print on the same figure, and mixing them answers a question about individual patients with an interval about an average.
Each subject was measured three times. Treating all 255 pairs as independent gives limits spanning 79.84 mmHg, while the method built for replicates gives 82.12 mmHg. What does treating replicates as independent actually cost?
Show the answer and why
Correct answer: The 95% confidence interval for each limit comes out as 4.30 mmHg either side, narrower than it should be, because the effective sample size is inflated
The limits themselves barely move, 79.84 against 82.12, because the two methods read the same subject at the same moment, so whatever the blood pressure happened to be cancels out of the difference. What collapses is the precision of the limits: the confidence interval goes from a correct 6.60 mmHg either side to 4.30 mmHg, because the data hold eighty-five independent people while the standard error was computed as though there were 255 independent observations. 74.22 mmHg is the genuinely too-narrow set, and it comes from a different approach, differencing each subject's mean of three readings with no correction, which describes a mean of three against a mean of three while every patient in clinic is measured once. A narrower interval claims more certainty than the data support, which pushes a device that failed towards passing.
Regressing the difference on the mean of the two methods gives a slope of 0.0356 mmHg/mmHg per subject, 95% CI -0.1028 to 0.1739, p = 0.610. What may be written?
Show the answer and why
Correct answer: The upper limit of 0.1739 multiplied by the hundred-plus mmHg range of these data is still not small, so all that may be written is that no trend was detected
Failing to reach significance usually means the sample was too small, not that the effect is zero. The upper limit of 0.1739 mmHg/mmHg across the hundred-plus mmHg range covered here leaves a tilt of a dozen or so mmHg between the ends of the range compatible with the data, and that is not a range small enough to declare absent. The correct sentence is that no trend of the difference with magnitude was detected, not that the difference is independent of magnitude; the second claim needs an equivalence margin fixed in advance. 0.6104 is the p value, and not detected is not the same as not there. 0.0514 is the slope on all the single-replicate pairs, and two flat estimates pointing the same way corroborate each other rather than failing.
Taking 140 mmHg as the cut-off, 24 of the 85 people are high by observer J and 36 are high by machine S. What does that say about replacing the observer with the machine?
Show the answer and why
Correct answer: Compared person by person, 14 cross the line purely because the measurement method changed
Comparing totals hides it. Person by person, 14 cross 140 mmHg purely because the method changed, 13 of them high on the machine and not on the observer and one the other way. 13 is that one-directional count, not the difference between the two totals. 24 is the number the observer called high, and whether the people the machine adds are truly hypertensive is a question these data cannot answer, because there is no third-party reference standard here; all they show is who changes side. Translating limits into clinical terms means asking whether decisions flip, not only reading the mean difference.
Same people, same afternoon, with only the identity of the other party swapped: two human observers against each other give a standard deviation of differences of just 2.26 mmHg on single readings. What does that control rule out?
Show the answer and why
Correct answer: Their limits span 8.86 mmHg, so the wide limits against the machine cannot be blamed on blood pressure being variable, on too few people, or on the method being harsh
Same people, same afternoon, same formulas, with only the identity of the other party changed. The two human observers span 8.86 mmHg on single readings, so the width against the machine is not blood pressure moving, not uncooperative subjects, and not a method that is harsh on everyone: the machine is what is wide. 30.63 mmHg is the standard deviation of the subject means, which describes how much the people differ from one another rather than any set of limits. 5.24 mmHg comes from differencing each person's mean of three readings, and averaging shrinks each method's own measurement error, so that figure describes a mean against a mean while every patient in clinic is measured once.
Chapters that use this method
Sources and licences
This page is original writing