Non-inferiority and equivalence
A result that is not statistically significant cannot be read backwards as a demonstration that the two treatments are comparable. Claiming comparability means writing down, before any data are seen, the largest loss you would still accept — the margin — and then asking which side of it the upper confidence limit falls on. This page puts one RCT interval in front of four different margins, watches the verdict travel from inferior to non-inferior, and shows that the tipping point sits exactly on the upper confidence limit.
“Not statistically significant” cannot be read backwards
The null hypothesis of a superiority trial is that the two arms do not differ. A test that fails to reject it says this study did not detect a difference — not that the difference is zero. An underpowered trial, a trial with too short a follow-up, a trial with a blunt outcome definition: each of them hands you an interval that crosses the null, and on paper those are indistinguishable from a trial of two genuinely comparable treatments.
Claiming comparability takes a different design. You have to write down a number before any data are seen: if the new treatment is worse than the old one, how much worse may it be before you would stop using it? That number is the margin. With it, the decision rule collapses to one sentence:
Take the difference, experimental minus reference, and ask which side of the margin the upper limit of its 95% confidence interval falls on.
Throughout this page a positive difference means worse. The margin is therefore a positive number sitting on the right of the plot, and an upper limit below the margin means that even the unfavourable end of the interval is still inside what was agreed to be acceptable.
The direction on this page is reversed, on purpose
The data are the ones the design chapter Randomised controlled trial already uses: medicaldata::indo_rct, 602 patients undergoing ERCP, with post-ERCP pancreatitis as the outcome. It is a superiority trial (Elmunzer BJ, et al. N Engl J Med 2012;366:1414-22, registration NCT00820612) and it never declared a margin of any kind.
Read in the direction the trial was actually run, the risk difference is indomethacin minus placebo: -7.8 percentage points, 95% confidence interval -13.1 to -2.5. The whole interval sits to the left of zero — on the better side. A non-inferiority margin sits on the worse side, so relative to this interval every positive margin lies to the right of it: all of them are cleared, and moving the margin cannot change the verdict. Demonstrating that a different margin gives a different answer is arithmetically impossible in this direction.
So this page turns the same data around and asks a different question: could the drug be left out? That is a question trials really do ask — de-implementation. Prophylaxis costs money, staff time and side effects, and if the price of omitting it is small enough it is worth saving. Asked that way, the experimental strategy becomes no prophylaxis, the reference becomes indomethacin, and the risk difference has the same magnitude with the opposite sign: 7.8 percentage points, 95% confidence interval 2.5 to 13.1. Now the margin lands near the interval, and moving it moves the verdict.
figures/scripts/B8-06-non-inferiority.ROne dataset, four margins, three verdicts
The four margins below were all invented for this page; nobody pre-specified them. Each is taken from one of the trial’s own design numbers so that they are at least defensible candidates rather than arbitrary picks. The verdicts, by contrast, are computed: the upper limit of the turned-around interval is compared with the margin, once.
| Margin | Where this margin comes from | Upper limit below the margin? | Verdict |
|---|---|---|---|
| 2 pp | concedes almost nothing — smaller than the difference this trial itself observed | no | inferior |
| 5 pp | the effect this trial was powered to detect (10% control risk down to 5%) | no | inconclusive |
| 10 pp | the whole control-arm risk this trial assumed at the design stage | no | inconclusive |
| 15 pp | deliberately lax, chosen so that the verdict flips | yes | non-inferior, not superior |
Same patients, same outcome, same interval, and the verdict travels from inferior all the way to non-inferior. The only thing exchanged is one number, and that number does not come out of the data.
The verdict changes between 10 and 15 percentage points. More precisely, it changes at 13.118 percentage points — which is the upper confidence limit itself. Any margin above it yields non-inferiority; any margin below it does not.
The widest margin: non-inferior, and statistically worse
The 15 pp row carries a second lesson of its own. Its verdict is non-inferiority, yet the lower limit of the turned-around interval is 2.5 percentage points, still above zero. Under that lax margin, omitting the drug is simultaneously:
- non-inferior — the unfavourable end of the interval (13.1 percentage points) is inside what was agreed to be acceptable
- statistically worse than the reference — the interval excludes zero altogether, so this study does detect a difference, and its direction is unfavourable
The two are not in conflict, because they answer different questions. Non-inferiority asks whether the loss is small enough, not whether there is one. Every row of the stats file carries a worse_than_reference flag, and all four margins have it set.
Equivalence is two-sided, and strictly harder
Non-inferiority polices one side: the experimental arm may not be much worse, and being much better is fine. Equivalence polices both — the entire interval must lie inside plus and minus the margin. Bioequivalence studies and method-comparison studies usually ask the second question.
| Margin | Non-inferiority verdict | Equivalence established? |
|---|---|---|
| plus and minus 2 pp | inferior | no |
| plus and minus 5 pp | inconclusive | no |
| plus and minus 10 pp | inconclusive | no |
| plus and minus 15 pp | non-inferior, not superior | yes |
Of the same four limits, only the widest establishes equivalence — and the way it does is worth noting: the lower limit 2.5 sits above minus 15 percentage points and the upper limit 13.1 sits below plus 15, so the whole interval is boxed in — even though the interval itself excludes zero. Equivalence bounds the magnitude, not the direction.
Five results a non-inferiority trial can report
The figure below is a diagram, not an analysis: its five intervals were chosen to illustrate the verdicts and are not computed from medicaldata::indo_rct or from any other data, so they must not be read as results. Their use is to let you place a published interval into one of five rows.
figures/scripts/B8-06-non-inferiority.R| Scenario | Design | Margin applies? | How to read it |
|---|---|---|---|
| superior (a superiority trial, no margin declared) | superiority | no | The whole interval lies on the better side of no difference. No margin was declared, so the margin line is not part of this trial's decision rule. |
| non-inferior and superior | non-inferiority | yes | The upper limit is below zero, so it is below the margin as well: both claims hold. |
| non-inferior, not superior | non-inferiority | yes | The upper limit is below the margin but the interval still contains zero. Non-inferiority is established; superiority is not, and must not be claimed. |
| inconclusive: the interval crosses the margin | non-inferiority | yes | The upper limit is beyond the margin. Nothing is established — this is the row that a paper must not report as 'comparable'. |
| inferior: worse by more than the margin | non-inferiority | yes | The whole interval lies beyond the margin: worse by more than the agreed amount. |
The fourth row needs the most care. An upper limit beyond the margin means nothing has been established: neither comparability nor inferiority. The honest word is inconclusive — not comparable, and certainly not “no difference shown, so the two are interchangeable”.
The margin is a clinical judgement, not a statistical choice
Because the margin decides the conclusion, its provenance has to be defensible. Convention requires that it is:
- Written into the protocol and fixed before recruitment. Choosing a margin after seeing the interval is painting the target around the arrow.
- Derived from the reference treatment’s historical effect against placebo. The usual construction allows the new treatment to give up only part of the old one’s benefit — a half, say — so that it still beats no treatment at all.
- Checked against the constancy assumption. That historical effect was measured in another era, another population, another standard of care. If today’s patients are at much lower risk, the old treatment’s effect shrinks too, and the margin has effectively widened on its own.
- Watched for biocreep. Each generation is non-inferior to the last, and far enough down the chain a treatment may be non-inferior to something that no longer works.
- Powered accordingly — usually a larger trial than a superiority version, because the margin is often smaller than the anticipated difference and squeezing an interval inside it takes a narrower interval. See Sample size and power.
Two intervals: Wald and Newcombe
This page uses the Wald interval for a difference of two independent proportions, because that is the one a non-inferiority report actually prints. The script also computes Newcombe’s hybrid-score interval, which combines the two Wilson intervals rather than assuming the difference itself is normal. Both in percentage points:
| Method | Lower limit | Upper limit |
|---|---|---|
| Wald (used throughout this page) | 2.45 | 13.12 |
| Newcombe hybrid-score | 2.40 | 13.16 |
They differ by less than a tenth of a percentage point, so at this number of events the choice changes no verdict at any of the four margins. That statement stops there; it is not a general one. When a risk sits close to 0 or 1, or when there are far fewer events, the normal approximation degrades and the two intervals separate — and the gap may be precisely the difference between clearing the margin and not. Non-inferiority designs run into this often, because both arms are frequently at low risk (the reference treatment works), so few events and risks near zero arrive together. The interval method has to be specified in advance, for the same reason the margin does: computing both afterwards and reporting the narrower one is choosing the result.
Intention-to-treat and per-protocol: this design reverses that too
An analysis that cannot be run can still have its concept quantified. The ladder below is a thought experiment on this trial’s own numbers: if some fraction of the experimental arm did not in fact follow the assigned strategy and behaved like the reference arm, the estimate shrinks toward zero by that fraction. The standard error is held fixed so that dilution is the only thing moving.
| Non-adherent fraction | Risk difference (pp) | 95% CI | Margin 10 pp | Margin 15 pp |
|---|---|---|---|---|
| 0% | 7.8 | 2.5 to 13.1 | inconclusive | non-inferior, not superior |
| 10% | 7.0 | 1.7 to 12.3 | inconclusive | non-inferior, not superior |
| 20% | 6.2 | 0.9 to 11.6 | inconclusive | non-inferior, not superior |
| 30% | 5.4 | 0.1 to 10.8 | inconclusive | non-inferior, not superior |
| 40% | 4.7 | -0.7 to 10.0 | inconclusive | non-inferior, not superior |
| 50% | 3.9 | -1.4 to 9.2 | non-inferior, not superior | non-inferior, not superior |
Look at the second column from the right. Under the 10 pp margin the verdict holds at inconclusive all the way to 40% non-adherence and only flips to non-inferiority at 50%. Dilution pushes the interval toward the safe side of the margin.
The intention-to-treat half is complete on this page: 307 patients assigned to no prophylaxis and 295 to indomethacin, with an observed outcome for all 602 records and nobody dropped for missing follow-up. What is missing is the other half, and this file cannot supply it.
What to look for when reading a paper
- What the margin is, and where it is stated. Can you find it in the protocol or the registry entry? If it appears only in the results, treat it as post hoc.
- What justifies that margin. Is the reference treatment’s historical effect against placebo cited, and what fraction of it is being preserved?
- The upper confidence limit. The verdict is that limit compared with the margin once, so reading the two together beats reading the concluding sentence.
- Whether intention-to-treat and per-protocol are both reported, and agree. A non-inferiority trial reporting only one deserves a discount.
- Whether superiority is claimed off the back of non-inferiority. That path has to be pre-specified in the testing sequence; the reverse direction is not available at all.
- The wording of the abstract. Not statistically significant, no difference detected, non-inferior and comparable are four different claims, and the commonest failure is an abstract that turns the first into the fourth.
Run it yourself
library(medicaldata)
data(indo_rct, package = "medicaldata")
d <- indo_rct
# Experimental strategy = omitting the drug; reference = indomethacin.
# A positive difference therefore means "omitting it is worse".
n1 <- sum(d$rx == "0_placebo")
e1 <- sum(d$rx == "0_placebo" & d$outcome == "1_yes")
n0 <- sum(d$rx == "1_indomethacin")
e0 <- sum(d$rx == "1_indomethacin" & d$outcome == "1_yes")
p1 <- e1 / n1
p0 <- e0 / n0
rd <- p1 - p0
se <- sqrt(p1 * (1 - p1) / n1 + p0 * (1 - p0) / n0)
ci <- rd + c(-1, 1) * qnorm(0.975) * se
round(100 * c(rd, ci), 2) # 7.79 2.45 13.12
# The entire decision rule, for a margin sitting on the worse side.
verdict <- function(lcl, ucl, margin) {
if (ucl < 0) "non-inferior and superior"
else if (ucl < margin) "non-inferior, not superior"
else if (lcl > margin) "inferior"
else "inconclusive"
}
for (m in c(0.02, 0.05, 0.10, 0.15)) {
cat(sprintf("margin %4.0f pp -> %s\n", 100 * m, verdict(ci[1], ci[2], m)))
}
# The margin at which the verdict changes IS the upper confidence limit.
100 * ci[2]
# Flip the sign and you have the trial as it was run: the interval then lies
# entirely below zero, so every positive margin is cleared and none can bite.
round(100 * (-rd + c(-1, 1) * qnorm(0.975) * se), 2)Verified with R 4.6.0 and medicaldata 0.2.0.
import numpy as np
from scipy.stats import norm
n1, e1 = 307, 52 # experimental strategy: no prophylaxis
n0, e0 = 295, 27 # reference: indomethacin
p1, p0 = e1 / n1, e0 / n0
rd = p1 - p0
se = np.sqrt(p1 * (1 - p1) / n1 + p0 * (1 - p0) / n0)
z = norm.ppf(0.975)
lcl, ucl = rd - z * se, rd + z * se
print(round(100 * rd, 2), round(100 * lcl, 2), round(100 * ucl, 2))
def verdict(lcl, ucl, margin):
if ucl < 0: return "non-inferior and superior"
if ucl < margin: return "non-inferior, not superior"
if lcl > margin: return "inferior"
return "inconclusive"
for m in (0.02, 0.05, 0.10, 0.15):
print(f"margin {100 * m:4.0f} pp -> {verdict(lcl, ucl, m)}")
print(100 * ucl) # the margin at which the verdict changesThere is no medicaldata package for Python, so the counts are entered directly below; they match the R output above. No special function is needed for the verdict — it is one comparison between the upper confidence limit and the margin.
Common misuses
| Misuse | Why it is wrong |
|---|---|
| A superiority test misses significance and the paper claims the treatments are comparable | That is a different design, and it needs a pre-specified margin; a difference not detected is not evidence of comparability |
| Setting the margin after seeing the interval | The margin decides the conclusion, so choosing it afterwards is painting the target around the arrow |
| Reporting “non-inferiority was established” as “the two work equally well” | Non-inferiority is a bounded claim; dropping the bound drops the only thing the design added |
| Reporting intention-to-treat alone in a non-inferiority trial | Dilution is anti-conservative in this direction; both analysis sets must be reported and must agree |
| Reusing the sample size of a superiority trial | The margin is often smaller than the anticipated difference, and squeezing the interval inside it takes a narrower interval |
| Claiming superiority off a non-inferiority result with no pre-specified testing sequence | That path has to be specified in advance; the reverse direction is not available at all |
| Writing an upper limit beyond the margin as “no difference shown, so interchangeable” | That is inconclusive: neither comparability nor inferiority has been established |
| Computing both interval methods afterwards and reporting the narrower | The same problem as choosing a margin afterwards; the method has to be pre-specified |
| A chain of treatments each non-inferior to the last | Biocreep: the newest may by now be non-inferior to something that does not work |
The page that one-line guard points at
Many places on this site stop the reader from turning “not statistically significant” into “the two are the same”, and each adds the same reminder: claiming comparability takes a margin fixed in advance, in a non-inferiority or an equivalence design (the pages differ on which, depending on whether the misreading they block is one-sided or two-sided — the t-test page says equivalence design). They run across all five families, from basic tests to evidence synthesis: t-tests and analysis of variance, the Cox proportional hazards model, the Brier score, NRI and IDI, causal mediation analysis and network meta-analysis.
A sentence can only deny the wrong reading. What this page adds is what the right one looks like: an interval, a line drawn in advance, and a claim that holds only once the two have been compared. Its home among the designs is Randomised controlled trial, whose intention-to-treat section already notes that non-inferiority reverses the direction; this page puts numbers on that remark.
Reproducing every number on this page
/opt/homebrew/bin/Rscript figures/scripts/B8-06-non-inferiority.RRead the figure
The answer comes from the same statistical output that produced this page's figures, not from a number typed in beside them.
One confidence interval set against four margins gives verdicts running from inferior all the way to non-inferior. Exactly which number does the verdict turn on?
Show the answer and why
Correct answer: 13.118 percentage points - not one of the tabulated margins at all, but the upper confidence limit the data themselves produced, which every wider margin clears and no narrower one does
13.118 is the upper limit of the reversed interval. A non-inferiority test is equivalent to asking whether the margin exceeds that upper limit, so the flip point is not any of the four margins in the table but the number the data themselves supply. 10.000 and 15.000 are the two margins the flip sits between, and a table can only tell you it happens somewhere in there; calling 15.000 the flip point merely reflects that it was the first margin tried, and a finer set of margins would have found the crossing earlier. The practical consequence: when a paper concludes that non-inferiority was met, half the information is in the data and the other half is in the protocol - you have to go back and ask who set that margin and when. That is why non-inferiority papers are reviewed against the protocol and the registration record rather than against the statistical methods section.
A thought experiment: suppose some fraction of the experimental arm did not follow the assigned strategy and behaved like the reference arm instead. Under the margin of 10 percentage points, how much non-adherence does it take before the verdict turns non-inferior?
Show the answer and why
Correct answer: 0.5. Dilution shrinks the point estimate towards zero and pushes the interval to the safe side of the margin, so the worse a trial is run the easier non-inferiority becomes - intention-to-treat is anti-conservative in this design
That column stays inconclusive all the way to 0.4 and turns non-inferior only at 0.5. The direction is the point: dilution makes the two arms look more alike, and looking more alike is precisely the conclusion a non-inferiority trial wants, so the worse the execution the easier the claim. A superiority trial is the mirror image - the same dilution makes a significant result harder to obtain, which is what makes intention-to-treat the conservative choice there. The verdict at 0.1 has not turned, so saying any non-adherence flips it overstates the sensitivity. 0.4 is the last row before the flip, and through that stretch dilution is moving the interval towards the safe side of the margin, not away from it. That is why per-protocol is not a secondary analysis in a non-inferiority design: both sets must be reported and must agree, and reporting intention-to-treat alone means concluding from the analysis that favours your own conclusion.
This page reverses the direction the trial was actually run in and asks instead whether the drug could be left out. Why not demonstrate margins in the trial's original direction? (The values below are absolute risk differences as proportions; multiply by a hundred for the percentage points shown on the page.)
Show the answer and why
Correct answer: Because the upper limit in the original direction is -0.0245, with the whole interval on the better side of zero, while a margin sits on the worse side - every positive margin clears, and moving it changes nothing
The upper limit in the original direction is -0.0245, with the whole interval to the left of zero, on the better side; a non-inferiority margin sits on the worse side, so every positive margin falls to the right of the interval, every one of them clears, and moving it changes no verdict. Demonstrating that a different margin gives a different conclusion is arithmetically impossible in that direction. 0.1312 is the upper limit of the reversed interval, the same magnitude with the opposite sign, because the two are one fact stated two ways. 0.0272 is the standard error of the difference, and it does set the width of the interval, but the issue was never whether the interval is wide enough - it is which side of the interval the margin falls on. The reversal changes no data, no population and no outcome definition, only the question, which matters because the alternative route - hunting for a subgroup with a wider interval - would amount to teaching post-hoc subgroup analysis.
Across the four margins the non-inferiority verdict runs from inferior to non-inferior. What do the equivalence verdicts do?
Show the answer and why
Correct answer: Only the margin of plus or minus 15 percentage points holds. Equivalence needs the whole interval inside both bounds at once, which is two-sided and far stricter
Only the widest margin holds: the lower limit has to sit above minus 15 percentage points and the upper limit below plus 15 before the whole interval is boxed in. Non-inferiority polices one side only - the experimental arm may not be much worse, and being much better is fine - so equivalence does not follow wherever non-inferiority holds, which is where the second option goes wrong. As for the strictest margin, an interval excluding zero says the data can detect a difference, which is a different question from whether the difference is small enough. In fact, under the widest margin, leaving the drug out is simultaneously non-inferior and statistically worse than the reference. Those do not conflict, because non-inferiority asks whether the difference is small enough, not whether there is one. Bioequivalence trials and method comparisons usually ask the two-sided question.
The script computes both a Wald and a Newcombe interval. The reversed Wald upper limit is 0.1312. What is the Newcombe upper limit, and what does it do to the verdicts? (Values are absolute risk differences as proportions.)
Show the answer and why
Correct answer: 0.1316. The two differ by less than a tenth of a percentage point, so at this event count neither choice changes any verdict - but that statement stops here, because with far fewer events the two intervals come apart
0.1316 differs from the Wald 0.1312 by less than a tenth of a percentage point, so at this event count both methods return the same verdicts. That is not a general rule: when one arm's risk sits close to zero or one, or when there are far fewer events, the normal approximation strains and the two intervals separate - and what separates them may be exactly whether the margin is crossed. Non-inferiority designs run into this often, because an effective reference treatment leaves both arms at low risk. 0.0240 is the Newcombe lower limit rather than the upper one, and reading a lower limit as an upper one produces the opposite conclusion that Newcombe is clearly narrower. As for -0.0240, that is the upper limit of the interval in the trial's original direction; the sign is reversed because the question was turned around, not because the method changed. The rule in practice is that the interval method is pre-specified, for the same reason the margin is: computing both after the fact and reporting the narrower one is choosing a result.
A non-inferiority trial is expected to report intention-to-treat and per-protocol together, and this page can produce only the first. Which statement is right?
Show the answer and why
Correct answer: The intention-to-treat set is complete: 307 were assigned to no prophylaxis and the rest to indomethacin, all with an observed outcome; what is missing is the other half, because this file records no adherence
307 is the number assigned to no prophylaxis, and all 602 rows of this file carry an observed outcome, so the intention-to-treat set has no gap. What is missing is per-protocol: the file holds baseline and procedural characteristics plus the randomised arm and the outcome, and no column records adherence, receipt of study drug, or protocol deviation. Using the event count of 52 to reconstruct adherence treats the outcome as an exposure, which destroys the randomisation outright, and not having an event is in any case a different thing from having complied. As for nobody being lost to follow-up so both sets must agree, loss to follow-up and non-adherence are two different gaps - somebody can be followed the whole way through and still not have taken the assigned strategy. The figure script therefore invents no proxy variable; it asserts instead that no such column exists, so that if the package ever adds one the script fails rather than continuing to claim a gap that has closed.
Each of the four margins on this page is drawn from one of the trial's own design numbers. The page says not one of them can stand as a defensible margin for this question. Why?
Show the answer and why
Correct answer: Because they were picked after the interval was known, and the widest of them, 15 percentage points, was chosen deliberately to make the verdict flip - a margin must be written into the protocol and fixed before recruitment, and setting it after seeing the interval is drawing the target round the arrow
The defect is in the timing, not in the values: all four were picked after the interval was known, and the one at 15 percentage points was chosen deliberately to make the verdict flip. Convention requires a margin to be written into the protocol and fixed before recruitment, because setting it after seeing the interval is drawing the target round the arrow - and a margin chosen afterwards looks exactly like a pre-specified one in the published paper. Saying a margin must exceed the observed difference inverts the rule: the strictest of the four fails not for being strict but for being retrospective, and a strict margin written down in advance is entirely legitimate. Saying only the historical effect is admissible overstates it too: deriving the margin from the reference treatment's effect against placebo, conceding only part of the older treatment's benefit so the new one still beats no treatment, is the conventional practice, but that margin fails here for the same reason of timing. Two further checks belong here: the constancy assumption, and biocreep - each generation is non-inferior to the last, and far enough down the chain the comparator may already be inert.
Chapters that use this method
Sources and licences
This page is original writing