Glossary

Statistical terms with their Chinese equivalents. The search box takes either language.

絕對風險下降absolute risk reductionARR
The difference in risk between groups, in percentage points. Reporting only the relative reduction overstates benefit, because the same RRR converts to a trivial ARR in a low-risk population.Appears in:D3-rct
Baujat plotBaujat plot
An influence diagnostic for meta-analysis, plotting each study's contribution to heterogeneity (Q) against its influence on the pooled estimate. Studies in the upper-right corner are the ones most in need of explanation — but being influential is never in itself a reason to exclude a study; the finding calls for looking at how its population, intervention or risk of bias differs.Appears in:B7-07-influence-diagnostics
Brier scoreBrier score
The mean squared error between predicted probabilities and observed outcomes. It mixes discrimination and calibration together, so a single value has no interpretable scale — it must be compared with a reference (such as predicting the population event rate for everyone), or read as a scaled Brier (Brier skill score). Computing it directly on censored survival data is wrong: restricting to those followed to the horizon silently changes the reference to the event rate among the followed-up, and the conclusion can reverse. The correct version is time-dependent and weighted by the censoring probability (IPCW).Appears in:B5-10-brier-nri-idiB4-04-roc-auc
校準calibration
Whether predicted probabilities match observed frequencies. A model with excellent discrimination can still be badly calibrated, and since clinical decisions use absolute probabilities, calibration often matters more.Appears in:B5-06-calibration
病例對照研究case-control study
Selects people by outcome and looks back at exposure. Because the case-to-control ratio is chosen by the investigator, incidence cannot be estimated — only odds ratios. Its strength is rare and slow diseases.Appears in:D2-case-control
設限censoring
The event has not occurred by the end of observation, so only a lower bound on the time is known. Discarding censored records or treating them as event-free both bias the estimate.Appears in:B3-01-censoring
Cohen's kappaCohen's kappaκ
Agreement between two raters after subtracting what chance alone would produce. The numerator is agreement beyond expectation; the denominator is how much agreement was left to achieve. Its least intuitive property is the kappa paradox: two tables with identical observed agreement can give kappas that differ several-fold, purely because the abnormal reading is rarer in one of them. The usual interpretive bands are convention, with no test behind them.Appears in:B9-01-kappaD4-diagnostic-accuracyD7-systematic-review
世代研究cohort study
Follows people forward from exposure to outcome. Because the denominator is known, incidence can be estimated directly. Prospective and retrospective differ in whether outcomes had occurred when the study began, not in direction.Appears in:D1-cohort
競爭風險competing risk
An event whose occurrence prevents the event of interest from ever happening. Under competing risks 1 − KM overestimates cumulative incidence, and the cumulative incidence function should be used instead.Appears in:B3-05-competing-risk
條件式邏輯迴歸conditional logistic regression
The correct analysis for matched case-control data: comparisons are made within matched sets and then pooled. Ordinary logistic regression breaks the matching, and the matching variables' own effects become inestimable.Appears in:D2-case-control
信賴區間confidence intervalCI
Under repeated sampling with the same method, 95% of such intervals contain the true value. It describes precision, not the probability that the truth lies inside this particular interval.
干擾confounding
A third variable influences both exposure and outcome, so the observed association is not the causal effect. Adjustment only handles measured confounders; unmeasured ones remain untouched in the estimate.Appears in:B6-01-dagD1-cohort
依適應症的干擾confounding by indication
Clinicians prescribe to patients who appear to need treatment, so receiving the drug also encodes being sicker. This is the hardest bias in observational drug research and the reason randomisation exists.Appears in:D1-cohortD3-rct
一致性consistency
The observed outcome under the treatment a person actually received equals their counterfactual outcome under that treatment. This requires the treatment to be well defined: exposures like exercise or weight loss have many versions, and when versions differ in effect the estimand has no single meaning.Appears in:B6-03-iptw
等高線漏斗圖contour-enhanced funnel plot
A funnel plot with the regions of statistical significance shaded. It has exactly one purpose, and it matters: distinguishing a gap that falls in the non-significant region (suggesting publication bias — the null studies were not published) from one that falls among small studies outside it (suggesting some other source of small-study effects, such as methodological quality or genuine heterogeneity). Asymmetry has five explanations; this plot rules some of them out.Appears in:B7-04-forest-funnel
累積統合分析cumulative meta-analysis
Recomputing the pooled estimate each time a study is added, in chronological order. Its most powerful use is showing the year by which the answer had stabilised, and how many more participants were recruited afterwards. But the procedure itself retests the same question every time a study arrives — the same phenomenon as multiplicity, in the time dimension — so a pooled estimate that crosses back over the null before settling is expected, not a sign of bad data. What addresses that repeated looking is a required information size and trial sequential analysis boundaries (see B7-06), not the cumulative plot itself.Appears in:B7-06-cumulative-meta
差異中的差異difference-in-differencesDiD
Subtracts the before-after change in an unaffected control series from the before-after change in the treated series. The counterfactual is borrowed rather than extrapolated, at the cost of assuming parallel trends.Appears in:B6-10-its-did
E 值E-value
The minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need with both exposure and outcome to fully explain away an observed association. Compute one for the point estimate and one for the confidence limit nearest the null; a large E-value means a strong confounder would be required, not that none exists.Appears in:B6-07-e-value
因療效提前中止early stopping for efficacy
Terminating a trial at an interim look under a pre-specified rule. The cost is systematic overestimation of the effect, because trials stop when their estimate happens to sit at the favourable end of its variation.Appears in:D3-rct
經驗校正empirical calibration
Uses the distribution of estimates from a set of negative controls, whose true effects are known to be null, as the empirical null. Any displacement from unity estimates the residual bias; calibrating against it widens confidence intervals and makes p-values more conservative.Appears in:B6-10-its-did
可交換性exchangeability
Conditional on the measured covariates, treatment assignment is independent of the counterfactual outcomes — that is, no unmeasured confounding. Randomisation buys it outright; observational studies can only assume conditional exchangeability, and the assumption is not testable from the data.Appears in:B6-03-iptwB6-06-target-trial
Fleiss kappaFleiss kappa
The agreement measure for three or more raters. It is not the average of the pairwise kappas; it is computed directly from how many raters assigned each subject to each category. Breaking it down by category is usually more informative than the single overall value, because disagreement tends to concentrate in one or two categories.Appears in:B9-01-kappa
Gray 檢定Gray's test
A test comparing cumulative incidence functions across groups under competing risks. The log-rank test compares cause-specific hazards and treats competing events as censoring, so under competing risks it can give a different — sometimes an opposite — conclusion. The two answer different questions (burden versus mechanism); neither is the wrong test.Appears in:B3-05-competing-risk
風險比hazard ratioHR
The ratio of the instantaneous event rate between two groups. HR = 1 means equal hazard. It is a ratio of rates, not a probability, and not a ratio of survival times.Appears in:B3-03-cox
異質性heterogeneity
Variation in true effects across studies. I² is a proportion, not a magnitude: it rises when studies are precise even if the differences are trivial. Use tau-squared or a prediction interval for absolute size.Appears in:B7-04-forest-funnel
Hodges-Lehmann 位移估計Hodges-Lehmann estimator
The effect size a non-parametric test should report: the median of all between-group pairwise differences, with a confidence interval, in the original units. Reporting only the Wilcoxon p-value says there is a difference without saying how large — and the size is what the clinic needs. It estimates a different quantity from log-transforming and comparing geometric means (a shift versus a ratio), and the two are interpreted quite differently.Appears in:B1-04-nonparametric
不死時間偏誤immortal time bias
When qualifying as exposed requires surviving to some later point, that guaranteed-event-free interval is credited to the exposed group's follow-up, manufacturing an apparent protective effect. Model exposure as time-varying instead.Appears in:B3-01-censoringD1-cohort
不一致性inconsistency
The degree to which direct and indirect evidence contradict each other in a network meta-analysis. ⚠️ A common trap: the inconsistency test a package prints by default may be the common-effect decomposition while the paper fits a random-effects model — on the same data one can be significant and the other show no signal at all. Check which model a quoted figure belongs to before citing it.Appears in:B7-05-network-meta
意向治療分析intention-to-treatITT
Analyse every participant in the group they were randomised to, regardless of the treatment actually received or adherence. It preserves the comparability randomisation bought, at the cost of a conservative effect estimate.Appears in:D3-rct
期中分析interim analysis
An analysis of accumulating data during a trial. Each look spends type I error, so an alpha-spending function allocates the total 0.05 across looks, leaving a final threshold slightly below 0.05.Appears in:D3-rct
中斷時間序列interrupted time seriesITS
Compares level and slope before and after an intervention in a single series, with the counterfactual extrapolated from the pre-intervention trend. No control group is needed, but secular trends and seasonality must be modelled or they are absorbed into the estimated effect.Appears in:B6-10-its-did
組內相關係數intraclass correlation coefficientICC
One variance decomposition, two uses. The clustering ICC asks how alike two people in the same cluster are; the reliability ICC asks how alike two measurements of the same subject are. The reliability side has six forms (single vs average measurement, consistency vs absolute agreement, raters random vs fixed), and a paper reporting "ICC = 0.85" without naming the form cannot be interpreted. Being a unitless proportion, it depends on how spread out the subjects are: the same instrument measured on a more homogeneous group gives a lower ICC.Appears in:B9-02-iccB2-07-clusteredB8-04-measurement-error
Kaplan-Meier 估計Kaplan-Meier estimatorKM
Estimates the survival function as a running product of conditional survival probabilities at each event time. The curve steps down only at events, and its right tail rests on ever fewer people.Appears in:B3-02-kaplan-meier
左截切left truncation
A person can only enter the study by surviving to some point, so before that point they are not part of the risk set. This differs from censoring: a censored person is in the data with incomplete time, whereas a left-truncated person should not be counted in the denominator at all before entry. Ignoring it systematically overstates survival.Appears in:B3-08-left-truncationB3-01-censoring
概似比likelihood ratioLR
The ratio of the probability of a test result in diseased versus non-diseased people. It converts pre-test odds into post-test odds directly and is unaffected by prevalence.Appears in:B4-03-likelihood-ratio
一致性界限limits of agreementLoA
What a Bland-Altman analysis produces: the mean difference between two measurement methods plus and minus 1.96 standard deviations of that difference, in the original units, so you can ask directly whether that width is clinically acceptable. It answers a different question from a correlation coefficient: a high r means the two move together, not that one can replace the other. When each subject is measured repeatedly the limits are computed differently, and treating the replicates as independent makes the confidence interval for the limits far too narrow.Appears in:B9-03-bland-altman
對數等級檢定log-rank test
Compares two entire survival curves. It yields a p-value but no effect size, and loses power badly when curves cross, because opposing differences cancel.Appears in:B3-02-kaplan-meier
邊際結構模型marginal structural modelMSM
A weighted model for time-varying confounding. When a covariate both confounds later treatment and is affected by earlier treatment, conditioning on it blocks a real pathway and can open a collider path; inverse-probability weighting instead builds a pseudo-population in which treatment is independent of the confounder.Appears in:B6-08-msm
中位存活時間median survival
The time at which the survival curve first reaches 0.5 — not the mean, which is usually much larger on a right-skewed distribution. When the curve never reaches 0.5 it is 'not reached', meaning beyond follow-up.Appears in:B3-02-kaplan-meier
統合分析meta-analysisMA
Statistically combining effect estimates across studies. It is a step inside a systematic review, not a substitute for one, and declining to pool is the right call when the studies are not clinically comparable.Appears in:B7-04-forest-funnelD7-systematic-review
特定時點存活率milestone survival
A comparison of survival probabilities at a pre-specified time point — for example, a 14 percentage point difference in three-year survival. It needs no proportional hazards assumption, but the time point must be fixed in advance, and a direct (Greenwood) interval and a log-log transformed interval do not agree.Appears in:B3-09-weighted-logrank
自然直接效果natural direct effectNDE
The effect of exposure on the outcome when the mediator is held at the value it would naturally take under no exposure. It sums with the natural indirect effect to give the total effect.Appears in:B6-09-mediation
自然間接效果natural indirect effectNIE
The part of the exposure effect transmitted through a change in the mediator. When its confidence interval crosses the null, report it as no detected mediation, not as evidence that the pathway is inactive.Appears in:B6-09-mediation
負對照negative control
An outcome that, by prior knowledge, should not be caused by the exposure, or an exposure that should not affect the outcome. If an association still appears, residual bias remains in the analysis.Appears in:B6-04-ivB6-10-its-did
淨重分類改善net reclassification improvementNRI
How many people a new marker moves into the "right" risk band. It has three structural problems: no null value to protect it; a poorly calibrated new model can still earn a positive NRI; and the categorical version depends entirely on where the cut-points are placed, so a different set gives a different conclusion. Most fundamentally it adds two proportions — one from the events, one from the non-events — when the two groups are usually of very different size, so a handful of future events harmed against a slightly larger handful of non-events helped can be reported as a healthy positive number. The question to ask is still net benefit.Appears in:B5-10-brier-nri-idi
網絡統合分析network meta-analysisNMA
A meta-analysis that pools several treatments at once, combining direct comparisons with indirect ones routed through a common comparator. As a result the league table prints pairs that have never been randomised against each other. The order of reading should be: the network plot, to see which pairs have direct evidence; then the pooled estimates; then whether direct and indirect agree (net-splitting); and only then the ranking.Appears in:B7-05-network-meta
nomogramnomogram
⚠️ This word names two different things on this site. A prediction-model nomogram (rms::nomogram) draws the linear predictor as a points scale: read off each variable, total the points, convert to a risk. It adds no information; the question to ask is whether the model behind it was internally validated, externally validated and calibrated. A Fagan nomogram is unrelated — a slide rule that multiplies pre-test odds by a likelihood ratio to give post-test probability, belonging to diagnostic reasoning.Appears in:B5-09-nomogramB4-03-likelihood-ratio
非劣性界值non-inferiority margin
The largest disadvantage a non-inferiority trial declares acceptable, specified in advance. The conclusion depends entirely on where it sits: the same data with a different margin gives a different answer, and where the margin came from is part of interpreting the trial, not a formality. Two things are routinely misread: establishing non-inferiority is not the same as showing equivalence (that needs an equivalence design); and a result can be both "non-inferior" and statistically worse than the reference. In a non-inferiority design the per-protocol analysis is not secondary — the dilution in ITT works in the favourable direction.Appears in:B8-06-non-inferiorityD3-rct
非資訊性設限non-informative censoring
Censoring is unrelated to the hazard at that moment. It holds at administrative end of study, but not when patients withdraw because they are deteriorating — and the data cannot reveal the difference.Appears in:B3-01-censoring
益一需治數number needed to treatNNT
The reciprocal of the absolute risk reduction: how many patients must be treated to prevent one event. It ties the effect size to baseline risk in a clinically legible unit.Appears in:D3-rct
勝算比odds ratioOR
The ratio of odds between groups. Case-control studies can only produce ORs because the denominators are chosen by the investigator. The OR approximates the RR only when the outcome is rare.Appears in:B2-02-logisticD2-case-control
平行趨勢parallel trends
The identifying assumption of difference-in-differences: absent the intervention, the two groups would have moved on parallel paths. It is an assumption about a counterfactual, so pre-intervention data can refute it but never confirm it.Appears in:B6-10-its-did
部分 AUCpartial AUCpAUC
The area under only a chosen stretch of the ROC curve, in specificity (or sensitivity). The reason to use it is that operation is confined to that stretch — screening and imaging studies typically work only at high specificity, and the whole-curve AUC counts a region that will never be used, so when two ROC curves cross the full AUC hides the difference. ⚠️ An unstandardised partial AUC cannot be compared across different stretches, because the maximum area differs; standardisation (McClish) maps it back linearly onto 0.5 to 1.Appears in:B4-05-cutoff
遵從方案分析per-protocol analysisPP
Analyse only participants who adhered to the protocol. It asks what happens under full adherence, but because exclusions occur after randomisation and may relate to prognosis, comparability is no longer guaranteed.Appears in:D3-rct
正性positivity
Every covariate pattern must have a non-zero probability of receiving each treatment. Violations show up clinically as extreme weights and poor overlap: with no comparable counterpart, the model can only extrapolate. Structural violations, where a group by definition cannot receive the treatment, cannot be fixed statistically.Appears in:B6-03-iptw
傾向分數propensity score
The conditional probability of receiving the exposure given measured covariates, used to match or weight so the groups balance. Balance is judged by standardised mean differences, not p-values, and only measured confounders are addressed.Appears in:B6-02-psm
中介比例proportion mediated
The share of the total effect carried by the indirect path. It becomes unstable when the total effect is near zero — the denominator vanishes, and the ratio can exceed one or turn negative — and should not be reported in that situation.Appears in:B6-09-mediation
比例風險假設proportional hazards assumptionPH
Cox regression assumes the hazard ratio between groups stays constant over follow-up. Crossing KM curves or a time trend in Schoenfeld residuals signal a violation, under which a single HR hides a time-varying effect.Appears in:B3-04-ph-assumption
發表偏誤publication bias
Non-significant studies are less likely to be published, tilting the visible literature toward positive findings. Funnel asymmetry is one signal, but heterogeneity, small-study quality, mathematical coupling and chance also produce it.Appears in:B7-04-forest-funnel
隨機效應模型random-effects model
Assumes each study estimates its own true effect drawn from a distribution, and pools the mean of that distribution. It differs from fixed-effect in its assumptions, not merely in conservatism.Appears in:B7-04-forest-funnel
隨機分配randomisation
Assignment decided by chance alone, making the arms equal in expectation on every characteristic including unmeasured ones — precisely what adjustment cannot do. It guarantees balance in expectation, not in any single allocation.Appears in:D3-rct
回憶偏誤recall bias
Cases search their memory harder than controls, producing more complete exposure reports and manufacturing an association. Mitigated by objective records and blinded interviewers.Appears in:D2-case-control
殘餘干擾residual confounding
Confounding that survives adjustment, from variables that were unmeasured, measured with error, or categorised too coarsely. No statistical test reveals it; only the limitations section does.Appears in:D1-cohort
限制平均存活時間restricted mean survival timeRMST
The area under the Kaplan-Meier curve between 0 and τ — on average, how long people lived during the first τ of follow-up. It needs no proportional hazards assumption and is reported in units of time, which makes it a substitute for the hazard ratio when proportionality clearly fails. The price is that τ must be fixed in advance, and the conclusion covers only what happens inside it.Appears in:B3-07-rmstB3-04-ph-assumption
相對風險risk ratioRR
The ratio of risks between groups. It must be read alongside the absolute risk reduction: relative measures are untethered from baseline risk and therefore always look larger.Appears in:D3-rct
ROC 曲線與曲線下面積ROC curve and AUC
Plots sensitivity against 1 − specificity across all thresholds; the area under it is the probability that a random diseased person scores higher than a random non-diseased one. It measures discrimination only, never calibration.Appears in:B4-04-roc-auc
選擇偏誤selection bias
Those included differ systematically from the target population in ways related to the research question. In case-control studies the classic form is controls drawn from a different source population than the cases.Appears in:D2-case-control
敏感度分析sensitivity analysis
Swapping out one analysis decision that could reasonably have gone the other way, rerunning, and asking whether the conclusion turns over. It varies analysis decisions, not data points (influence analysis), populations (subgroup analysis) or significance thresholds (multiplicity adjustment) — and so needs no multiplicity adjustment, because it asks about stability rather than significance. Stating that a result was robust carries no information on its own: what belongs in the report is which perturbations were run, the effect estimate under each, and whether any turned the conclusion over. It rules out only what it perturbed; a bias shared by every setting survives all of them.Appears in:B8-07-sensitivity-analysis
敏感度與特異度sensitivity and specificity
Sensitivity is the proportion of diseased correctly identified; specificity the proportion of non-diseased correctly excluded. Both are properties of the test and do not vary with prevalence — but predictive values do.Appears in:B4-01-sens-spec
標準化平均差standardised mean differenceSMD
The between-group difference divided by the pooled standard deviation. Unlike a p-value it does not grow with sample size, which is why balance checks use |SMD| < 0.1 rather than p > 0.05.Appears in:B1-01-table1B6-02-psm
統計顯著性statistical significance
A non-significant result means this study did not detect a difference — not that the groups are equivalent. Claiming equivalence requires a non-inferiority design with a pre-specified margin.
隨機優勢stochastic ordering
What the Wilcoxon rank-sum (Mann-Whitney) test actually tests: whether P(X > Y) equals 1/2 — the probability that a randomly drawn member of group A exceeds a randomly drawn member of group B. It does not test medians. So "the medians are identical but the test is significant" is not a contradiction; a difference in distributional shape alone can move P(X > Y) away from 1/2.Appears in:B1-04-nonparametric
SUCRA/P-scoreSUCRA / P-score
A number between 0 and 1 that compresses a treatment's ranking probabilities in a network meta-analysis. It carries no information about effect size: the top-ranked treatment may beat the runner-up by a hair, with a head-to-head interval that crosses zero. The uncertainty in a ranking is usually far larger than the ranking itself suggests, so it should not be read as "which drug is best", nor trusted before inconsistency has been checked.Appears in:B7-05-network-meta
綜合 ROC 曲線summary ROC curveSROC
The central plot of a diagnostic accuracy meta-analysis: each study's sensitivity and specificity as a point, with a summary curve through them. It exists because studies use different thresholds, which makes sensitivity and specificity negatively correlated across studies, so pooling the two separately misstates the joint uncertainty. The confidence ellipse is the uncertainty about where the summary point lies; the prediction ellipse is where the true value of the next study might fall. The two are routinely confused, and can differ by an order of magnitude in size.Appears in:B7-08-diagnostic-metaD4-diagnostic-accuracy
系統性回顧systematic reviewSR
A reproducible search and appraisal of all studies meeting pre-specified criteria. The review is the process; meta-analysis is an optional step within it, and a meta-analysis without one is a red flag.Appears in:D7-systematic-review
隨時間變動的 AUCtime-dependent AUCAUC(t)
A survival model's discrimination at a particular time t, taken as the area under the time-dependent ROC curve. Harrell's C compresses discrimination across all time points into one number, whereas clinical decisions are made at particular time points; estimating it requires inverse probability of censoring weighting to handle censored observations.Appears in:B5-08-time-dependent-aucB5-05-external-validation
時間相依變項time-varying covariate
A covariate whose value changes during follow-up. The data are expanded into (start, stop] intervals with several rows per person; this is also the correct fix for immortal time bias.Appears in:B3-06-time-varying
trim-and-filltrim-and-fill
After detecting funnel-plot asymmetry, an algorithm "fills in" mirrored studies on the empty side and recomputes the pooled value. The added studies are generated, not recovered, so this is a sensitivity analysis — an answer to how far the pooled estimate could move — not a correction to it. No method can retrieve studies that were never published.Appears in:B7-04-forest-funnel
小提琴圖violin plot
A distribution shown as a mirrored kernel density, usually with a box or the individual points (raincloud) laid over it. It exists as the counterexample to the bar-with-error-bar (dynamite) plot, which hides the sample size, the shape, any bimodality and the outliers — and whose error bar may be an SD, an SE or a 95% CI. Those three differ substantially in length and not at all in appearance, so the caption has to say which it is.Appears in:B1-05-distribution-plotsB1-02-t-test-anova
加權 kappaweighted kappa
A kappa that penalises a one-category disagreement less than a three-category one, for ordinal rating scales (BI-RADS, Gleason, tumour grade). Weights are usually linear or quadratic. Applied to nominal categories it is arithmetic rather than analysis: there is no distance between the categories, and the partial credit inflates expected agreement faster than observed, so the weighted value can come out below the unweighted one.Appears in:B9-01-kappa
加權 log-rank 檢定weighted log-rank test
A log-rank statistic in which each event time's observed-minus-expected contribution is multiplied by a time-varying weight. The Fleming-Harrington family uses two parameters, ρ and γ, to tilt the weight towards early or late follow-up; the delayed effects seen with immunotherapy call for weights that favour later times. The weighting must be pre-specified.Appears in:B3-09-weighted-logrankB3-02-kaplan-meier