Glossary
Statistical terms with their Chinese equivalents. The search box takes either language.
No matching terms.
- 絕對風險下降absolute risk reductionARR
- The difference in risk between groups, in percentage points. Reporting only the relative reduction overstates benefit, because the same RRR converts to a trivial ARR in a low-risk population.Appears in:D3-rct
- Baujat plotBaujat plot
- An influence diagnostic for meta-analysis, plotting each study's contribution to heterogeneity (Q) against its influence on the pooled estimate. Studies in the upper-right corner are the ones most in need of explanation — but being influential is never in itself a reason to exclude a study; the finding calls for looking at how its population, intervention or risk of bias differs.Appears in:B7-07-influence-diagnostics
- Brier scoreBrier score
- The mean squared error between predicted probabilities and observed outcomes. It mixes discrimination and calibration together, so a single value has no interpretable scale — it must be compared with a reference (such as predicting the population event rate for everyone), or read as a scaled Brier (Brier skill score). Computing it directly on censored survival data is wrong: restricting to those followed to the horizon silently changes the reference to the event rate among the followed-up, and the conclusion can reverse. The correct version is time-dependent and weighted by the censoring probability (IPCW).Appears in:B5-10-brier-nri-idiB4-04-roc-auc
- 校準calibration
- Whether predicted probabilities match observed frequencies. A model with excellent discrimination can still be badly calibrated, and since clinical decisions use absolute probabilities, calibration often matters more.Appears in:B5-06-calibration
- 病例對照研究case-control study
- Selects people by outcome and looks back at exposure. Because the case-to-control ratio is chosen by the investigator, incidence cannot be estimated — only odds ratios. Its strength is rare and slow diseases.Appears in:D2-case-control
- 設限censoring
- The event has not occurred by the end of observation, so only a lower bound on the time is known. Discarding censored records or treating them as event-free both bias the estimate.Appears in:B3-01-censoring
- Cohen's kappaCohen's kappaκ
- Agreement between two raters after subtracting what chance alone would produce. The numerator is agreement beyond expectation; the denominator is how much agreement was left to achieve. Its least intuitive property is the kappa paradox: two tables with identical observed agreement can give kappas that differ several-fold, purely because the abnormal reading is rarer in one of them. The usual interpretive bands are convention, with no test behind them.Appears in:B9-01-kappaD4-diagnostic-accuracyD7-systematic-review
- 世代研究cohort study
- Follows people forward from exposure to outcome. Because the denominator is known, incidence can be estimated directly. Prospective and retrospective differ in whether outcomes had occurred when the study began, not in direction.Appears in:D1-cohort
- 競爭風險competing risk
- An event whose occurrence prevents the event of interest from ever happening. Under competing risks 1 − KM overestimates cumulative incidence, and the cumulative incidence function should be used instead.Appears in:B3-05-competing-risk
- 條件式邏輯迴歸conditional logistic regression
- The correct analysis for matched case-control data: comparisons are made within matched sets and then pooled. Ordinary logistic regression breaks the matching, and the matching variables' own effects become inestimable.Appears in:D2-case-control
- 信賴區間confidence intervalCI
- Under repeated sampling with the same method, 95% of such intervals contain the true value. It describes precision, not the probability that the truth lies inside this particular interval.
- 干擾confounding
- A third variable influences both exposure and outcome, so the observed association is not the causal effect. Adjustment only handles measured confounders; unmeasured ones remain untouched in the estimate.Appears in:B6-01-dagD1-cohort
- 依適應症的干擾confounding by indication
- Clinicians prescribe to patients who appear to need treatment, so receiving the drug also encodes being sicker. This is the hardest bias in observational drug research and the reason randomisation exists.Appears in:D1-cohortD3-rct
- 一致性consistency
- The observed outcome under the treatment a person actually received equals their counterfactual outcome under that treatment. This requires the treatment to be well defined: exposures like exercise or weight loss have many versions, and when versions differ in effect the estimand has no single meaning.Appears in:B6-03-iptw
- 等高線漏斗圖contour-enhanced funnel plot
- A funnel plot with the regions of statistical significance shaded. It has exactly one purpose, and it matters: distinguishing a gap that falls in the non-significant region (suggesting publication bias — the null studies were not published) from one that falls among small studies outside it (suggesting some other source of small-study effects, such as methodological quality or genuine heterogeneity). Asymmetry has five explanations; this plot rules some of them out.Appears in:B7-04-forest-funnel
- 累積統合分析cumulative meta-analysis
- Recomputing the pooled estimate each time a study is added, in chronological order. Its most powerful use is showing the year by which the answer had stabilised, and how many more participants were recruited afterwards. But the procedure itself retests the same question every time a study arrives — the same phenomenon as multiplicity, in the time dimension — so a pooled estimate that crosses back over the null before settling is expected, not a sign of bad data. What addresses that repeated looking is a required information size and trial sequential analysis boundaries (see B7-06), not the cumulative plot itself.Appears in:B7-06-cumulative-meta
- 差異中的差異difference-in-differencesDiD
- Subtracts the before-after change in an unaffected control series from the before-after change in the treated series. The counterfactual is borrowed rather than extrapolated, at the cost of assuming parallel trends.Appears in:B6-10-its-did
- E 值E-value
- The minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need with both exposure and outcome to fully explain away an observed association. Compute one for the point estimate and one for the confidence limit nearest the null; a large E-value means a strong confounder would be required, not that none exists.Appears in:B6-07-e-value
- 因療效提前中止early stopping for efficacy
- Terminating a trial at an interim look under a pre-specified rule. The cost is systematic overestimation of the effect, because trials stop when their estimate happens to sit at the favourable end of its variation.Appears in:D3-rct
- 經驗校正empirical calibration
- Uses the distribution of estimates from a set of negative controls, whose true effects are known to be null, as the empirical null. Any displacement from unity estimates the residual bias; calibrating against it widens confidence intervals and makes p-values more conservative.Appears in:B6-10-its-did
- 可交換性exchangeability
- Conditional on the measured covariates, treatment assignment is independent of the counterfactual outcomes — that is, no unmeasured confounding. Randomisation buys it outright; observational studies can only assume conditional exchangeability, and the assumption is not testable from the data.Appears in:B6-03-iptwB6-06-target-trial
- Fleiss kappaFleiss kappa
- The agreement measure for three or more raters. It is not the average of the pairwise kappas; it is computed directly from how many raters assigned each subject to each category. Breaking it down by category is usually more informative than the single overall value, because disagreement tends to concentrate in one or two categories.Appears in:B9-01-kappa
- Gray 檢定Gray's test
- A test comparing cumulative incidence functions across groups under competing risks. The log-rank test compares cause-specific hazards and treats competing events as censoring, so under competing risks it can give a different — sometimes an opposite — conclusion. The two answer different questions (burden versus mechanism); neither is the wrong test.Appears in:B3-05-competing-risk
- 風險比hazard ratioHR
- The ratio of the instantaneous event rate between two groups. HR = 1 means equal hazard. It is a ratio of rates, not a probability, and not a ratio of survival times.Appears in:B3-03-cox
- 異質性heterogeneity
- Variation in true effects across studies. I² is a proportion, not a magnitude: it rises when studies are precise even if the differences are trivial. Use tau-squared or a prediction interval for absolute size.Appears in:B7-04-forest-funnel
- Hodges-Lehmann 位移估計Hodges-Lehmann estimator
- The effect size a non-parametric test should report: the median of all between-group pairwise differences, with a confidence interval, in the original units. Reporting only the Wilcoxon p-value says there is a difference without saying how large — and the size is what the clinic needs. It estimates a different quantity from log-transforming and comparing geometric means (a shift versus a ratio), and the two are interpreted quite differently.Appears in:B1-04-nonparametric
- 不死時間偏誤immortal time bias
- When qualifying as exposed requires surviving to some later point, that guaranteed-event-free interval is credited to the exposed group's follow-up, manufacturing an apparent protective effect. Model exposure as time-varying instead.Appears in:B3-01-censoringD1-cohort
- 不一致性inconsistency
- The degree to which direct and indirect evidence contradict each other in a network meta-analysis. ⚠️ A common trap: the inconsistency test a package prints by default may be the common-effect decomposition while the paper fits a random-effects model — on the same data one can be significant and the other show no signal at all. Check which model a quoted figure belongs to before citing it.Appears in:B7-05-network-meta
- 意向治療分析intention-to-treatITT
- Analyse every participant in the group they were randomised to, regardless of the treatment actually received or adherence. It preserves the comparability randomisation bought, at the cost of a conservative effect estimate.Appears in:D3-rct
- 期中分析interim analysis
- An analysis of accumulating data during a trial. Each look spends type I error, so an alpha-spending function allocates the total 0.05 across looks, leaving a final threshold slightly below 0.05.Appears in:D3-rct
- 中斷時間序列interrupted time seriesITS
- Compares level and slope before and after an intervention in a single series, with the counterfactual extrapolated from the pre-intervention trend. No control group is needed, but secular trends and seasonality must be modelled or they are absorbed into the estimated effect.Appears in:B6-10-its-did
- 組內相關係數intraclass correlation coefficientICC
- One variance decomposition, two uses. The clustering ICC asks how alike two people in the same cluster are; the reliability ICC asks how alike two measurements of the same subject are. The reliability side has six forms (single vs average measurement, consistency vs absolute agreement, raters random vs fixed), and a paper reporting "ICC = 0.85" without naming the form cannot be interpreted. Being a unitless proportion, it depends on how spread out the subjects are: the same instrument measured on a more homogeneous group gives a lower ICC.Appears in:B9-02-iccB2-07-clusteredB8-04-measurement-error
- Kaplan-Meier 估計Kaplan-Meier estimatorKM
- Estimates the survival function as a running product of conditional survival probabilities at each event time. The curve steps down only at events, and its right tail rests on ever fewer people.Appears in:B3-02-kaplan-meier
- 左截切left truncation
- A person can only enter the study by surviving to some point, so before that point they are not part of the risk set. This differs from censoring: a censored person is in the data with incomplete time, whereas a left-truncated person should not be counted in the denominator at all before entry. Ignoring it systematically overstates survival.Appears in:B3-08-left-truncationB3-01-censoring
- 概似比likelihood ratioLR
- The ratio of the probability of a test result in diseased versus non-diseased people. It converts pre-test odds into post-test odds directly and is unaffected by prevalence.Appears in:B4-03-likelihood-ratio
- 一致性界限limits of agreementLoA
- What a Bland-Altman analysis produces: the mean difference between two measurement methods plus and minus 1.96 standard deviations of that difference, in the original units, so you can ask directly whether that width is clinically acceptable. It answers a different question from a correlation coefficient: a high r means the two move together, not that one can replace the other. When each subject is measured repeatedly the limits are computed differently, and treating the replicates as independent makes the confidence interval for the limits far too narrow.Appears in:B9-03-bland-altman
- 對數等級檢定log-rank test
- Compares two entire survival curves. It yields a p-value but no effect size, and loses power badly when curves cross, because opposing differences cancel.Appears in:B3-02-kaplan-meier
- 邊際結構模型marginal structural modelMSM
- A weighted model for time-varying confounding. When a covariate both confounds later treatment and is affected by earlier treatment, conditioning on it blocks a real pathway and can open a collider path; inverse-probability weighting instead builds a pseudo-population in which treatment is independent of the confounder.Appears in:B6-08-msm
- 中位存活時間median survival
- The time at which the survival curve first reaches 0.5 — not the mean, which is usually much larger on a right-skewed distribution. When the curve never reaches 0.5 it is 'not reached', meaning beyond follow-up.Appears in:B3-02-kaplan-meier
- 統合分析meta-analysisMA
- Statistically combining effect estimates across studies. It is a step inside a systematic review, not a substitute for one, and declining to pool is the right call when the studies are not clinically comparable.Appears in:B7-04-forest-funnelD7-systematic-review
- 特定時點存活率milestone survival
- A comparison of survival probabilities at a pre-specified time point — for example, a 14 percentage point difference in three-year survival. It needs no proportional hazards assumption, but the time point must be fixed in advance, and a direct (Greenwood) interval and a log-log transformed interval do not agree.Appears in:B3-09-weighted-logrank
- 自然直接效果natural direct effectNDE
- The effect of exposure on the outcome when the mediator is held at the value it would naturally take under no exposure. It sums with the natural indirect effect to give the total effect.Appears in:B6-09-mediation
- 自然間接效果natural indirect effectNIE
- The part of the exposure effect transmitted through a change in the mediator. When its confidence interval crosses the null, report it as no detected mediation, not as evidence that the pathway is inactive.Appears in:B6-09-mediation
- 負對照negative control
- An outcome that, by prior knowledge, should not be caused by the exposure, or an exposure that should not affect the outcome. If an association still appears, residual bias remains in the analysis.Appears in:B6-04-ivB6-10-its-did
- 淨重分類改善net reclassification improvementNRI
- How many people a new marker moves into the "right" risk band. It has three structural problems: no null value to protect it; a poorly calibrated new model can still earn a positive NRI; and the categorical version depends entirely on where the cut-points are placed, so a different set gives a different conclusion. Most fundamentally it adds two proportions — one from the events, one from the non-events — when the two groups are usually of very different size, so a handful of future events harmed against a slightly larger handful of non-events helped can be reported as a healthy positive number. The question to ask is still net benefit.Appears in:B5-10-brier-nri-idi
- 網絡統合分析network meta-analysisNMA
- A meta-analysis that pools several treatments at once, combining direct comparisons with indirect ones routed through a common comparator. As a result the league table prints pairs that have never been randomised against each other. The order of reading should be: the network plot, to see which pairs have direct evidence; then the pooled estimates; then whether direct and indirect agree (net-splitting); and only then the ranking.Appears in:B7-05-network-meta
- nomogramnomogram
- ⚠️ This word names two different things on this site. A prediction-model nomogram (rms::nomogram) draws the linear predictor as a points scale: read off each variable, total the points, convert to a risk. It adds no information; the question to ask is whether the model behind it was internally validated, externally validated and calibrated. A Fagan nomogram is unrelated — a slide rule that multiplies pre-test odds by a likelihood ratio to give post-test probability, belonging to diagnostic reasoning.Appears in:B5-09-nomogramB4-03-likelihood-ratio
- 非劣性界值non-inferiority margin
- The largest disadvantage a non-inferiority trial declares acceptable, specified in advance. The conclusion depends entirely on where it sits: the same data with a different margin gives a different answer, and where the margin came from is part of interpreting the trial, not a formality. Two things are routinely misread: establishing non-inferiority is not the same as showing equivalence (that needs an equivalence design); and a result can be both "non-inferior" and statistically worse than the reference. In a non-inferiority design the per-protocol analysis is not secondary — the dilution in ITT works in the favourable direction.Appears in:B8-06-non-inferiorityD3-rct
- 非資訊性設限non-informative censoring
- Censoring is unrelated to the hazard at that moment. It holds at administrative end of study, but not when patients withdraw because they are deteriorating — and the data cannot reveal the difference.Appears in:B3-01-censoring
- 益一需治數number needed to treatNNT
- The reciprocal of the absolute risk reduction: how many patients must be treated to prevent one event. It ties the effect size to baseline risk in a clinically legible unit.Appears in:D3-rct
- 勝算比odds ratioOR
- The ratio of odds between groups. Case-control studies can only produce ORs because the denominators are chosen by the investigator. The OR approximates the RR only when the outcome is rare.Appears in:B2-02-logisticD2-case-control
- 平行趨勢parallel trends
- The identifying assumption of difference-in-differences: absent the intervention, the two groups would have moved on parallel paths. It is an assumption about a counterfactual, so pre-intervention data can refute it but never confirm it.Appears in:B6-10-its-did
- 部分 AUCpartial AUCpAUC
- The area under only a chosen stretch of the ROC curve, in specificity (or sensitivity). The reason to use it is that operation is confined to that stretch — screening and imaging studies typically work only at high specificity, and the whole-curve AUC counts a region that will never be used, so when two ROC curves cross the full AUC hides the difference. ⚠️ An unstandardised partial AUC cannot be compared across different stretches, because the maximum area differs; standardisation (McClish) maps it back linearly onto 0.5 to 1.Appears in:B4-05-cutoff
- 遵從方案分析per-protocol analysisPP
- Analyse only participants who adhered to the protocol. It asks what happens under full adherence, but because exclusions occur after randomisation and may relate to prognosis, comparability is no longer guaranteed.Appears in:D3-rct
- 正性positivity
- Every covariate pattern must have a non-zero probability of receiving each treatment. Violations show up clinically as extreme weights and poor overlap: with no comparable counterpart, the model can only extrapolate. Structural violations, where a group by definition cannot receive the treatment, cannot be fixed statistically.Appears in:B6-03-iptw
- 傾向分數propensity score
- The conditional probability of receiving the exposure given measured covariates, used to match or weight so the groups balance. Balance is judged by standardised mean differences, not p-values, and only measured confounders are addressed.Appears in:B6-02-psm
- 中介比例proportion mediated
- The share of the total effect carried by the indirect path. It becomes unstable when the total effect is near zero — the denominator vanishes, and the ratio can exceed one or turn negative — and should not be reported in that situation.Appears in:B6-09-mediation
- 比例風險假設proportional hazards assumptionPH
- Cox regression assumes the hazard ratio between groups stays constant over follow-up. Crossing KM curves or a time trend in Schoenfeld residuals signal a violation, under which a single HR hides a time-varying effect.Appears in:B3-04-ph-assumption
- 發表偏誤publication bias
- Non-significant studies are less likely to be published, tilting the visible literature toward positive findings. Funnel asymmetry is one signal, but heterogeneity, small-study quality, mathematical coupling and chance also produce it.Appears in:B7-04-forest-funnel
- 隨機效應模型random-effects model
- Assumes each study estimates its own true effect drawn from a distribution, and pools the mean of that distribution. It differs from fixed-effect in its assumptions, not merely in conservatism.Appears in:B7-04-forest-funnel
- 隨機分配randomisation
- Assignment decided by chance alone, making the arms equal in expectation on every characteristic including unmeasured ones — precisely what adjustment cannot do. It guarantees balance in expectation, not in any single allocation.Appears in:D3-rct
- 回憶偏誤recall bias
- Cases search their memory harder than controls, producing more complete exposure reports and manufacturing an association. Mitigated by objective records and blinded interviewers.Appears in:D2-case-control
- 殘餘干擾residual confounding
- Confounding that survives adjustment, from variables that were unmeasured, measured with error, or categorised too coarsely. No statistical test reveals it; only the limitations section does.Appears in:D1-cohort
- 限制平均存活時間restricted mean survival timeRMST
- The area under the Kaplan-Meier curve between 0 and τ — on average, how long people lived during the first τ of follow-up. It needs no proportional hazards assumption and is reported in units of time, which makes it a substitute for the hazard ratio when proportionality clearly fails. The price is that τ must be fixed in advance, and the conclusion covers only what happens inside it.Appears in:B3-07-rmstB3-04-ph-assumption
- 相對風險risk ratioRR
- The ratio of risks between groups. It must be read alongside the absolute risk reduction: relative measures are untethered from baseline risk and therefore always look larger.Appears in:D3-rct
- ROC 曲線與曲線下面積ROC curve and AUC
- Plots sensitivity against 1 − specificity across all thresholds; the area under it is the probability that a random diseased person scores higher than a random non-diseased one. It measures discrimination only, never calibration.Appears in:B4-04-roc-auc
- 選擇偏誤selection bias
- Those included differ systematically from the target population in ways related to the research question. In case-control studies the classic form is controls drawn from a different source population than the cases.Appears in:D2-case-control
- 敏感度分析sensitivity analysis
- Swapping out one analysis decision that could reasonably have gone the other way, rerunning, and asking whether the conclusion turns over. It varies analysis decisions, not data points (influence analysis), populations (subgroup analysis) or significance thresholds (multiplicity adjustment) — and so needs no multiplicity adjustment, because it asks about stability rather than significance. Stating that a result was robust carries no information on its own: what belongs in the report is which perturbations were run, the effect estimate under each, and whether any turned the conclusion over. It rules out only what it perturbed; a bias shared by every setting survives all of them.Appears in:B8-07-sensitivity-analysis
- 敏感度與特異度sensitivity and specificity
- Sensitivity is the proportion of diseased correctly identified; specificity the proportion of non-diseased correctly excluded. Both are properties of the test and do not vary with prevalence — but predictive values do.Appears in:B4-01-sens-spec
- 標準化平均差standardised mean differenceSMD
- The between-group difference divided by the pooled standard deviation. Unlike a p-value it does not grow with sample size, which is why balance checks use |SMD| < 0.1 rather than p > 0.05.Appears in:B1-01-table1B6-02-psm
- 統計顯著性statistical significance
- A non-significant result means this study did not detect a difference — not that the groups are equivalent. Claiming equivalence requires a non-inferiority design with a pre-specified margin.
- 隨機優勢stochastic ordering
- What the Wilcoxon rank-sum (Mann-Whitney) test actually tests: whether P(X > Y) equals 1/2 — the probability that a randomly drawn member of group A exceeds a randomly drawn member of group B. It does not test medians. So "the medians are identical but the test is significant" is not a contradiction; a difference in distributional shape alone can move P(X > Y) away from 1/2.Appears in:B1-04-nonparametric
- SUCRA/P-scoreSUCRA / P-score
- A number between 0 and 1 that compresses a treatment's ranking probabilities in a network meta-analysis. It carries no information about effect size: the top-ranked treatment may beat the runner-up by a hair, with a head-to-head interval that crosses zero. The uncertainty in a ranking is usually far larger than the ranking itself suggests, so it should not be read as "which drug is best", nor trusted before inconsistency has been checked.Appears in:B7-05-network-meta
- 綜合 ROC 曲線summary ROC curveSROC
- The central plot of a diagnostic accuracy meta-analysis: each study's sensitivity and specificity as a point, with a summary curve through them. It exists because studies use different thresholds, which makes sensitivity and specificity negatively correlated across studies, so pooling the two separately misstates the joint uncertainty. The confidence ellipse is the uncertainty about where the summary point lies; the prediction ellipse is where the true value of the next study might fall. The two are routinely confused, and can differ by an order of magnitude in size.Appears in:B7-08-diagnostic-metaD4-diagnostic-accuracy
- 系統性回顧systematic reviewSR
- A reproducible search and appraisal of all studies meeting pre-specified criteria. The review is the process; meta-analysis is an optional step within it, and a meta-analysis without one is a red flag.Appears in:D7-systematic-review
- 隨時間變動的 AUCtime-dependent AUCAUC(t)
- A survival model's discrimination at a particular time t, taken as the area under the time-dependent ROC curve. Harrell's C compresses discrimination across all time points into one number, whereas clinical decisions are made at particular time points; estimating it requires inverse probability of censoring weighting to handle censored observations.Appears in:B5-08-time-dependent-aucB5-05-external-validation
- 時間相依變項time-varying covariate
- A covariate whose value changes during follow-up. The data are expanded into (start, stop] intervals with several rows per person; this is also the correct fix for immortal time bias.Appears in:B3-06-time-varying
- trim-and-filltrim-and-fill
- After detecting funnel-plot asymmetry, an algorithm "fills in" mirrored studies on the empty side and recomputes the pooled value. The added studies are generated, not recovered, so this is a sensitivity analysis — an answer to how far the pooled estimate could move — not a correction to it. No method can retrieve studies that were never published.Appears in:B7-04-forest-funnel
- 小提琴圖violin plot
- A distribution shown as a mirrored kernel density, usually with a box or the individual points (raincloud) laid over it. It exists as the counterexample to the bar-with-error-bar (dynamite) plot, which hides the sample size, the shape, any bimodality and the outliers — and whose error bar may be an SD, an SE or a 95% CI. Those three differ substantially in length and not at all in appearance, so the caption has to say which it is.Appears in:B1-05-distribution-plotsB1-02-t-test-anova
- 加權 kappaweighted kappa
- A kappa that penalises a one-category disagreement less than a three-category one, for ordinal rating scales (BI-RADS, Gleason, tumour grade). Weights are usually linear or quadratic. Applied to nominal categories it is arithmetic rather than analysis: there is no distance between the categories, and the partial credit inflates expected agreement faster than observed, so the weighted value can come out below the unweighted one.Appears in:B9-01-kappa
- 加權 log-rank 檢定weighted log-rank test
- A log-rank statistic in which each event time's observed-minus-expected contribution is multiplied by a time-varying weight. The Fleming-Harrington family uses two parameters, ρ and γ, to tilt the weight towards early or late follow-up; the delayed effects seen with immunotherapy call for weights that favour later times. The weighting must be pre-specified.Appears in:B3-09-weighted-logrankB3-02-kaplan-meier