參考文獻
共 206 筆,依主題分組。每筆書目都以 PubMed、Crossref、arXiv 或 GitHub 實際查證過; 標「AI 章 [N]」者為第 12 關內文的引用編號。各章節內文另附該章使用的文獻連結。
總論與報告規範(15 筆)
- Moher D, Cook DJ, Eastwood S et al. Improving the quality of reports of meta-analyses of randomised controlled trials: the QUOROM statement. Quality of Reporting of Meta-analyses. Lancet. 1999.
- Stroup DF, Berlin JA, Morton SC et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. Meta-analysis Of Observational Studies in Epidemiology (MOOSE) group. JAMA. 2000.
- Booth A, Clarke M, Ghersi D et al. An international registry of systematic-review protocols. Lancet. 2011.
- Hutton B, Salanti G, Caldwell DM et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions: checklist and explanations. Ann Intern Med. 2015.
- Shamseer L, Moher D, Clarke M et al. Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-P) 2015: elaboration and explanation. BMJ. 2015.
- McGowan J, Sampson M, Salzwedel DM et al. PRESS Peer Review of Electronic Search Strategies: 2015 Guideline Statement. J Clin Epidemiol. 2016.
- Whiting P, Savović J, Higgins JP et al. ROBIS: A new tool to assess risk of bias in systematic reviews was developed. J Clin Epidemiol. 2016.
- Page MJ, Moher D Evaluations of the uptake and impact of the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) Statement and extensions: a scoping review. Syst Rev. 2017.
- McInnes MDF, Moher D, Thombs BD et al. Preferred Reporting Items for a Systematic Review and Meta-analysis of Diagnostic Test Accuracy Studies: The PRISMA-DTA Statement. JAMA. 2018.
- Cumpston M, Li T, Page MJ et al. Updated guidance for trusted systematic reviews: a new edition of the Cochrane Handbook for Systematic Reviews of Interventions. Cochrane Database Syst Rev. 2019.
- Higgins JPT, Thomas J, Chandler J, et al. (eds) Cochrane Handbook for Systematic Reviews of Interventions (version 6; living online version 6.5). Wiley / Cochrane (book). 2019.
- Page MJ, McKenzie JE, Bossuyt PM et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021.
- Page MJ, Moher D, Bossuyt PM et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021.
- Rethlefsen ML, Kirtley S, Waffenschmidt S et al. PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Syst Rev. 2021.
- Veroniki AA, Hutton B, Stevens A et al. Update to the PRISMA guidelines for network meta-analyses and scoping reviews and development of guidelines for rapid reviews: a scoping review protocol. JBI Evid Synth. 2025.
問題形成與搜尋(5 筆)
- McAuley L, Pham B, Tugwell P et al. Does the inclusion of grey literature influence estimates of intervention effectiveness reported in meta-analyses?. Lancet. 2000.
- Hopewell S, McDonald S, Clarke M et al. Grey literature in meta-analyses of randomized trials of health care interventions. Cochrane Database Syst Rev. 2007.
- Methley AM, Campbell S, Chew-Graham C et al. PICO, PICOS and SPIDER: a comparison study of specificity and sensitivity in three search tools for qualitative systematic reviews. BMC Health Serv Res. 2014.
- Bramer WM, Rethlefsen ML, Kleijnen J et al. Optimal database combinations for literature searches in systematic reviews: a prospective exploratory study. Syst Rev. 2017.
- Paez A Gray literature: An important resource in systematic reviews. J Evid Based Med. 2017.
篩選與萃取(10 筆)
- Landis JR, Koch GG The measurement of observer agreement for categorical data. Biometrics. 1977.
- Edwards P, Clarke M, DiGuiseppi C et al. Identification of randomized controlled trials in systematic reviews: accuracy and reliability of screening records. Stat Med. 2002.
- Elbourne DR, Altman DG, Higgins JP et al. Meta-analyses involving cross-over trials: methodological issues. Int J Epidemiol. 2002.
- Hozo SP, Djulbegovic B, Hozo I. Estimating the mean and variance from the median, range, and the size of a sample. BMC Med Res Methodol. 2005.
- Buscemi N, Hartling L, Vandermeer B et al. Single data extraction generated more errors than double data extraction in systematic reviews. J Clin Epidemiol. 2006.
- Altman DG, Bland JM. How to obtain the confidence interval from a P value. BMJ. 2011.
- McHugh ML Interrater reliability: the kappa statistic. Biochem Med (Zagreb). 2012.
- Wan X, Wang W, Liu J et al. Estimating the sample mean and standard deviation from the sample size, median, range and/or interquartile range. BMC Med Res Methodol. 2014.
- Luo D, Wan X, Liu J et al. Optimally estimating the sample mean from the sample size, median, mid-range, and/or mid-quartile range. Stat Methods Med Res. 2018.
- Waffenschmidt S, Knelangen M, Sieben W et al. Single screening versus conventional double screening for study selection in systematic reviews: a methodological systematic review. BMC Med Res Methodol. 2019.
Risk of bias(偏誤風險)評估(14 筆)
- Jüni P, Witschi A, Bloch R et al. The hazards of scoring the quality of clinical trials for meta-analysis. JAMA. 1999.
- Jüni P, Altman DG, Egger M Systematic reviews in health care: Assessing the quality of controlled clinical trials. BMJ. 2001.
- Hartling L, Ospina M, Liang Y et al. Risk of bias versus quality assessment of randomised controlled trials: cross sectional study. BMJ. 2009.
- Stang A Critical evaluation of the Newcastle-Ottawa scale for the assessment of the quality of nonrandomized studies in meta-analyses. Eur J Epidemiol. 2010.
- Higgins JP, Altman DG, Gøtzsche PC et al. The Cochrane Collaboration's tool for assessing risk of bias in randomised trials. BMJ. 2011.
- Whiting PF, Rutjes AW, Westwood ME et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011.
- Savović J, Jones HE, Altman DG et al. Influence of reported study design characteristics on intervention effect estimates from randomized, controlled trials. Ann Intern Med. 2012.
- Sterne JA, Hernán MA, Reeves BC et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016.
- Shea BJ, Reeves BC, Wells G et al. AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both. BMJ. 2017.
- Savovic J, Turner RM, Mawdsley D et al. Association Between Risk-of-Bias Assessments and Results of Randomized Trials in Cochrane Reviews: The ROBES Meta-Epidemiologic Study. Am J Epidemiol. 2018.
- Sterne JAC, Savović J, Page MJ et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019.
- Minozzi S, Dwan K, Borrelli F et al. Reliability of the revised Cochrane risk-of-bias tool for randomised trials (RoB2) improved with the use of implementation instruction. J Clin Epidemiol. 2022.
- Page MJ, Sterne JAC, Boutron I et al. ROB-ME: a tool for assessing risk of bias due to missing evidence in systematic reviews with meta-analysis. BMJ. 2023.
- Higgins JPT, Morgan RL, Rooney AA et al. A tool to assess risk of bias in non-randomized follow-up studies of exposure effects (ROBINS-E). Environ Int. 2024.
效果量與合併模型(25 筆)
- MANTEL N, HAENSZEL W Statistical aspects of the analysis of data from retrospective studies of disease. J Natl Cancer Inst. 1959.
- Greenland S, Robins JM Estimation of a common effect parameter from sparse follow-up data. Biometrics. 1985.
- DerSimonian R, Laird N Meta-analysis in clinical trials. Control Clin Trials. 1986.
- Colditz GA, Brewer TF, Berkey CS et al. Efficacy of BCG vaccine in the prevention of tuberculosis. Meta-analysis of the published literature. JAMA. 1994.
- Hackshaw AK, Law MR, Wald NJ. The accumulated evidence on lung cancer and environmental tobacco smoke. BMJ. 1997.
- Davies HT, Crombie IK, Tavakoli M. When can odds ratios mislead?. BMJ. 1998.
- Zhang J, Yu KF. What's the relative risk? A method of correcting the odds ratio in cohort studies of common outcomes. JAMA. 1998.
- Normand SL. Meta-analysis: formulating, evaluating, combining, and reporting. Stat Med. 1999.
- Chinn S. A simple method for converting an odds ratio to effect size for use in meta-analysis. Stat Med. 2000.
- Hartung J, Knapp G A refined method for the meta-analysis of controlled clinical trials with binary outcome. Stat Med. 2001.
- Altman DG, Deeks JJ. Meta-analysis, Simpson's paradox, and the number needed to treat. BMC Med Res Methodol. 2002.
- Deeks JJ. Issues in the selection of a summary statistic for meta-analysis of clinical trials with binary outcomes. Stat Med. 2002.
- Sidik K, Jonkman JN A simple confidence interval for meta-analysis. Stat Med. 2002.
- Knapp G, Hartung J Improved tests for a random effects meta-regression with a single covariate. Stat Med. 2003.
- Sweeting MJ, Sutton AJ, Lambert PC What to add to nothing? Use and avoidance of continuity corrections in meta-analysis of sparse data. Stat Med. 2004.
- Bradburn MJ, Deeks JJ, Berlin JA et al. Much ado about nothing: a comparison of the performance of meta-analytical methods with rare events. Stat Med. 2007.
- Tierney JF, Stewart LA, Ghersi D et al. Practical methods for incorporating summary time-to-event data into meta-analysis. Trials. 2007.
- Higgins JP, Thompson SG, Spiegelhalter DJ A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009.
- Borenstein M, Hedges LV, Higgins JP et al. A basic introduction to fixed-effect and random-effects models for meta-analysis. Res Synth Methods. 2010.
- Riley RD, Higgins JP, Deeks JJ Interpretation of random effects meta-analyses. BMJ. 2011.
- IntHout J, Ioannidis JP, Borm GF The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med Res Methodol. 2014.
- IntHout J, Ioannidis JP, Rovers MM et al. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016.
- Veroniki AA, Jackson D, Viechtbauer W et al. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Res Synth Methods. 2016.
- Page MJ, Altman DG, McKenzie JE et al. Flaws in the application and interpretation of statistical analyses in systematic reviews of therapeutic interventions were common: a cross-sectional analysis. J Clin Epidemiol. 2018.
- Langan D, Higgins JPT, Jackson D et al. A comparison of heterogeneity variance estimators in simulated random-effects meta-analyses. Res Synth Methods. 2019.
異質性(10 筆)
- Thompson SG, Sharp SJ Explaining heterogeneity in meta-analysis: a comparison of methods. Stat Med. 1999.
- Higgins JP, Thompson SG Quantifying heterogeneity in a meta-analysis. Stat Med. 2002.
- Thompson SG, Higgins JP How should meta-regression analyses be undertaken and interpreted?. Stat Med. 2002.
- Higgins JP, Thompson SG, Deeks JJ et al. Measuring inconsistency in meta-analyses. BMJ. 2003.
- Viechtbauer W Confidence intervals for the amount of heterogeneity in meta-analysis. Stat Med. 2007.
- Higgins JP Commentary: Heterogeneity in meta-analysis should be expected and appropriately quantified. Int J Epidemiol. 2008.
- Rücker G, Schwarzer G, Carpenter JR et al. Undue reliance on I(2) in assessing heterogeneity may mislead. BMC Med Res Methodol. 2008.
- Viechtbauer W, Cheung MW Outlier and influence diagnostics for meta-analysis. Res Synth Methods. 2010.
- von Hippel PT The heterogeneity statistic I(2) can be biased in small meta-analyses. BMC Med Res Methodol. 2015.
- Borenstein M, Higgins JP, Hedges LV et al. Basics of meta-analysis: I(2) is not an absolute measure of heterogeneity. Res Synth Methods. 2017.
Publication bias 與 small-study effects(14 筆)
- Dickersin K The existence of publication bias and risk factors for its occurrence. JAMA. 1990.
- Begg CB, Mazumdar M Operating characteristics of a rank correlation test for publication bias. Biometrics. 1994.
- Egger M, Davey Smith G, Schneider M et al. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997.
- Duval S, Tweedie R Trim and fill: A simple funnel-plot-based method of testing and adjusting for publication bias in meta-analysis. Biometrics. 2000.
- Sutton AJ, Duval SJ, Tweedie RL et al. Empirical assessment of effect of publication bias on meta-analyses. BMJ. 2000.
- Terrin N, Schmid CH, Lau J et al. Adjusting for publication bias in the presence of heterogeneity. Stat Med. 2003.
- Lau J, Ioannidis JP, Terrin N et al. The case of the misleading funnel plot. BMJ. 2006.
- Peters JL, Sutton AJ, Jones DR et al. Comparison of two methods to detect publication bias in meta-analysis. JAMA. 2006.
- Ioannidis JP, Trikalinos TA The appropriateness of asymmetry tests for publication bias in meta-analyses: a large survey. CMAJ. 2007.
- Peters JL, Sutton AJ, Jones DR et al. Contour-enhanced meta-analysis funnel plots help distinguish publication bias from other causes of asymmetry. J Clin Epidemiol. 2008.
- Rücker G, Schwarzer G, Carpenter J Arcsine test for publication bias in meta-analyses with binary outcomes. Stat Med. 2008.
- Sterne JA, Sutton AJ, Ioannidis JP et al. Recommendations for examining and interpreting funnel plot asymmetry in meta-analyses of randomised controlled trials. BMJ. 2011.
- Lin L, Chu H Quantifying publication bias in meta-analysis. Biometrics. 2018.
- Page MJ, Sterne JAC, Higgins JPT et al. Investigating and dealing with publication bias and other reporting biases in meta-analyses of health research: A review. Res Synth Methods. 2021.
Certainty of evidence:GRADE(15 筆)
- Guyatt GH, Oxman AD, Vist GE et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008.
- Balshem H, Helfand M, Schünemann HJ et al. GRADE guidelines: 3. Rating the quality of evidence. J Clin Epidemiol. 2011.
- Guyatt G, Oxman AD, Akl EA et al. GRADE guidelines: 1. Introduction-GRADE evidence profiles and summary of findings tables. J Clin Epidemiol. 2011.
- Guyatt GH, Oxman AD, Kunz R et al. GRADE guidelines 6. Rating the quality of evidence--imprecision. J Clin Epidemiol. 2011.
- Guyatt GH, Oxman AD, Kunz R et al. GRADE guidelines: 2. Framing the question and deciding on important outcomes. J Clin Epidemiol. 2011.
- Guyatt GH, Oxman AD, Kunz R et al. GRADE guidelines: 7. Rating the quality of evidence--inconsistency. J Clin Epidemiol. 2011.
- Guyatt GH, Oxman AD, Kunz R et al. GRADE guidelines: 8. Rating the quality of evidence--indirectness. J Clin Epidemiol. 2011.
- Guyatt GH, Oxman AD, Montori V et al. GRADE guidelines: 5. Rating the quality of evidence--publication bias. J Clin Epidemiol. 2011.
- Guyatt GH, Oxman AD, Sultan S et al. GRADE guidelines: 9. Rating up the quality of evidence. J Clin Epidemiol. 2011.
- Guyatt GH, Oxman AD, Vist G et al. GRADE guidelines: 4. Rating the quality of evidence--study limitations (risk of bias). J Clin Epidemiol. 2011.
- Brunetti M, Shemilt I, Pregno S et al. GRADE guidelines: 10. Considering resource use and rating the quality of economic evidence. J Clin Epidemiol. 2013.
- Guyatt GH, Oxman AD, Santesso N et al. GRADE guidelines: 12. Preparing summary of findings tables-binary outcomes. J Clin Epidemiol. 2013.
- Guyatt GH, Thorlund K, Oxman AD et al. GRADE guidelines: 13. Preparing summary of findings tables and evidence profiles-continuous outcomes. J Clin Epidemiol. 2013.
- Schünemann HJ, Cuello C, Akl EA et al. GRADE guidelines: 18. How ROBINS-I and other tools to assess risk of bias in nonrandomized studies should be used to rate the certainty of a body of evidence. J Clin Epidemiol. 2019.
- Zhang Y, Coello PA, Guyatt GH et al. GRADE guidelines: 20. Assessing the certainty of evidence in the importance of outcomes or values and preferences-inconsistency, imprecision, and other domains. J Clin Epidemiol. 2019.
進階主題:NMA、DTA、IPD、TSA、Bayesian(26 筆)
- Antman EM, Lau J, Kupelnick B et al. A comparison of results of meta-analyses of randomized control trials and recommendations of clinical experts. Treatments for myocardial infarction. JAMA. 1992.
- Lau J, Antman EM, Jimenez-Silva J et al. Cumulative meta-analysis of therapeutic trials for myocardial infarction. N Engl J Med. 1992.
- Hasselblad V. Meta-analysis of multitreatment studies. Med Decis Making. 1998.
- Rutter CM, Gatsonis CA A hierarchical regression approach to meta-analysis of diagnostic test accuracy evaluations. Stat Med. 2001.
- Sutton AJ, Abrams KR Bayesian methods in meta-analysis and evidence synthesis. Stat Methods Med Res. 2001.
- Lu G, Ades AE Combination of direct and indirect evidence in mixed treatment comparisons. Stat Med. 2004.
- Deeks JJ, Macaskill P, Irwig L The performance of tests of publication bias and other sample size effects in systematic reviews of diagnostic test accuracy was assessed. J Clin Epidemiol. 2005.
- Reitsma JB, Glas AS, Rutjes AW et al. Bivariate analysis of sensitivity and specificity produces informative summary measures in diagnostic reviews. J Clin Epidemiol. 2005.
- Leeflang MM, Deeks JJ, Gatsonis C et al. Systematic reviews of diagnostic test accuracy. Ann Intern Med. 2008.
- Salanti G, Higgins JP, Ades AE et al. Evaluation of networks of randomized trials. Stat Methods Med Res. 2008.
- Wetterslev J, Thorlund K, Brok J et al. Trial sequential analysis may establish when firm evidence is reached in cumulative meta-analysis. J Clin Epidemiol. 2008.
- Brok J, Thorlund K, Wetterslev J et al. Apparently conclusive meta-analyses may be inconclusive--Trial sequential analysis adjustment of random error risk due to repetitive testing of accumulating data in apparently conclusive neonatal meta-analyses. Int J Epidemiol. 2009.
- Wetterslev J, Thorlund K, Brok J et al. Estimating required information size by quantifying diversity in random-effects model meta-analyses. BMC Med Res Methodol. 2009.
- Riley RD, Lambert PC, Abo-Zaid G Meta-analysis of individual participant data: rationale, conduct, and reporting. BMJ. 2010.
- Salanti G, Ades AE, Ioannidis JP Graphical methods and numerical summaries for presenting results from multiple-treatment meta-analysis: an overview and tutorial. J Clin Epidemiol. 2011.
- Higgins JP, Jackson D, Barrett JK et al. Consistency and inconsistency in network meta-analysis: concepts and models for multi-arm studies. Res Synth Methods. 2012.
- Salanti G Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: many names, many benefits, many concerns for the next generation evidence synthesis tool. Res Synth Methods. 2012.
- Chaimani A, Higgins JP, Mavridis D et al. Graphical tools for network meta-analysis in STATA. PLoS One. 2013.
- Dias S, Welton NJ, Sutton AJ et al. Evidence synthesis for decision making 4: inconsistency in networks of evidence based on randomized controlled trials. Med Decis Making. 2013.
- Salanti G, Del Giovane C, Chaimani A et al. Evaluating the quality of evidence from a network meta-analysis. PLoS One. 2014.
- Rücker G, Schwarzer G Ranking treatments in frequentist network meta-analysis works without resampling methods. BMC Med Res Methodol. 2015.
- Stewart LA, Clarke M, Rovers M et al. Preferred Reporting Items for Systematic Review and Meta-Analyses of individual participant data: the PRISMA-IPD Statement. JAMA. 2015.
- Imberger G, Thorlund K, Gluud C et al. False-positive findings in Cochrane meta-analyses with and without application of trial sequential analysis: an empirical review. BMJ Open. 2016.
- Rouse B, Chaimani A, Li T Network meta-analysis: an introduction for clinicians. Intern Emerg Med. 2017.
- Cipriani A, Furukawa TA, Salanti G et al. Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: a systematic review and network meta-analysis. Lancet. 2018.
- Nikolakopoulou A, Higgins JPT, Papakonstantinou T et al. CINeMA: An approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020.
經典警世案例(12 筆)
- Teo KK, Yusuf S, Collins R et al. Effects of intravenous magnesium in suspected acute myocardial infarction: overview of randomised trials. BMJ. 1991.
- Woods KL, Fletcher S, Roffe C et al. Intravenous magnesium sulphate in suspected acute myocardial infarction: results of the second Leicester Intravenous Magnesium Intervention Trial (LIMIT-2). Lancet. 1992.
- Yusuf S, Teo K, Woods K Intravenous magnesium in acute myocardial infarction. An effective, safe, simple, and inexpensive intervention. Circulation. 1993.
- (group author) ISIS-4: a randomised factorial trial assessing early oral captopril, oral mononitrate, and intravenous magnesium sulphate in 58,050 patients with suspected acute myocardial infarction. ISIS-4 (Fourth International Study of Infarct Survival) Collaborative Group. Lancet. 1995.
- LeLorier J, Grégoire G, Benhaddad A et al. Discrepancies between meta-analyses and subsequent large randomized, controlled trials. N Engl J Med. 1997.
- Ioannidis JP Contradicted and initially stronger effects in highly cited clinical research. JAMA. 2005.
- Nissen SE, Wolski K Effect of rosiglitazone on the risk of myocardial infarction and death from cardiovascular causes. N Engl J Med. 2007.
- Psaty BM, Furberg CD Rosiglitazone and cardiovascular risk. N Engl J Med. 2007.
- Home PD, Pocock SJ, Beck-Nielsen H et al. Rosiglitazone evaluated for cardiovascular outcomes in oral agent combination therapy for type 2 diabetes (RECORD): a multicentre, randomised, open-label trial. Lancet. 2009.
- Ioannidis JP Meta-research: The art of getting it wrong. Res Synth Methods. 2010.
- Ioannidis JP The Mass Production of Redundant, Misleading, and Conflicted Systematic Reviews and Meta-analyses. Milbank Q. 2016.
- Page MJ, Shamseer L, Altman DG et al. Epidemiology and Reporting Characteristics of Systematic Reviews of Biomedical Research: A Cross-Sectional Study. PLoS Med. 2016.
LLM / AI 在 SR/MA 的實證(35 筆)
- Bhattacharyya M, Miller VM, Bhattacharyya D et al. High Rates of Fabricated and Inaccurate References in ChatGPT-Generated Medical Content. Cureus. 2023.
- Walters WH, Wilder EI Fabrication and errors in the bibliographic citations generated by ChatGPT. Sci Rep. 2023.
- Chelli M, Descamps J, Lavoué V et al. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. J Med Internet Res. 2024.
- Dennstädt F, Zink J, Putora PM et al. Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain. Syst Rev. 2024.
- Gartlehner G, Kahwati L, Hilscher R et al. Data extraction for evidence synthesis using a large language model: A proof-of-concept study. Res Synth Methods. 2024.
- Guo E, Gupta M, Deng J et al. Automated Paper Screening for Clinical Reviews Using Large Language Models: Data Analysis Study. J Med Internet Res. 2024.
- Hasan B, Saadi S, Rajjoub NS et al. Integrating large language models in systematic reviews: a framework and case study using ROBINS-I for risk of bias assessment. BMJ Evid Based Med. 2024.
- Khraisha Q, Put S, Kappenberg J et al. Can large language models replace humans in systematic reviews? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages. Res Synth Methods. 2024.
- Lai H, Ge L, Sun M et al. Assessing the Risk of Bias in Randomized Clinical Trials With Large Language Models. JAMA Netw Open. 2024.
- Matsui K, Utsumi T, Aoki Y et al. Human-Comparable Sensitivity of Large Language Models in Identifying Eligible Studies Through Title and Abstract Screening: 3-Layer Strategy Using GPT-3.5 and GPT-4 for Systematic Reviews. J Med Internet Res. 2024.
- Oami T, Okada Y, Nakada TA Performance of a Large Language Model in Screening Citations. JAMA Netw Open. 2024.
- Tran VT, Gartlehner G, Yaacoub S et al. Sensitivity and Specificity of Using GPT-3.5 Turbo Models for Title and Abstract Screening in Systematic Reviews and Meta-analyses. Ann Intern Med. 2024.
- Cao C, Arora R, Cento P, et al. Automation of Systematic Reviews with Large Language Models (otto-SR). medRxiv (preprint; not peer reviewed at time of search). 2025.
- Cao C, Sang J, Arora R et al. Development of Prompt Templates for Large Language Model-Driven Screening in Systematic Reviews. Ann Intern Med. 2025.
- Eisele-Metzger A, Lieberum JL, Toews M et al. Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2. Res Synth Methods. 2025.
- Flemyng E, Noel-Storr A, Macura B et al. Position Statement on Artificial Intelligence (AI) Use in Evidence Synthesis Across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence 2025. Campbell Syst Rev. 2025.
- Flemyng E, Noel-Storr A, Macura B et al. Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence 2025. JBI Evid Synth. 2025.
- Gartlehner G, Kugley S, Crotty K et al. Artificial Intelligence-Assisted Data Extraction With a Large Language Model: A Study Within Reviews. Ann Intern Med. 2025.
- Gartlehner G, Nussbaumer-Streit B, Hamel C et al. Responsible Integration of Artificial Intelligence in Rapid Reviews: A Position Statement From the Cochrane Rapid Reviews Methods Group. Cochrane Evid Synth Methods. 2025.
- Huang J, Lai H, Zhao W et al. Large Language Model–Assisted Risk-of-Bias Assessment in Randomized Controlled Trials Using the Revised Risk-of-Bias Tool: Evaluation Study. J Med Internet Res. 2025.
- Lieberum JL, Toews M, Metzendorf MI et al. Large language models for conducting systematic reviews: on the rise, but not yet ready for use-a scoping review. J Clin Epidemiol. 2025.
- RAISE working group (Cochrane, Campbell Collaboration, JBI, Collaboration for Environmental Evidence; J Thomas among leads) Responsible use of AI in evidence SynthEsis (RAISE) 1-3: recommendations and guidance. OSF collection (web resource). 2025.
- Rose CJ, Bidonde J, Ringsten M et al. Using a Large Language Model (ChatGPT-4o) to Assess the Risk of Bias in Randomized Controlled Trials of Medical Interventions: Interrater Agreement With Human Reviewers. Cochrane Evid Synth Methods. 2025.
- Scherbakov D, Hubig N, Jansari V et al. The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review. J Am Med Inform Assoc. 2025.
- Siemens W, von Elm E, Binder H et al. Opportunities, challenges and risks of using artificial intelligence for evidence synthesis. BMJ Evid Based Med. 2025.
- Beber SA, Groff KD, Mange TR et al. Not Ready for Prime Time: Limitations of a Retrieval-Augmented Generation Large Language Model in Assessing Risk of Bias in Observational Studies. J Pediatr Soc North Am. 2026.
- Gandhi S, Shokravi A, Chelliahpillai Y et al. Evaluating large language model performance in Risk of Bias assessments: A cross-sectional validation study. PLoS One. 2026.
- Gartlehner G, Banda S, Callaghan M et al. Cochrane evaluation of (semi-)automated review methods: protocol for an adaptive platform study within reviews. J Clin Epidemiol. 2026.
- Lai YJ, Lin SH, Liu JW Evaluating the Accuracy of Large Language Models in Risk-of-Bias Assessment Using Version 2 of the Cochrane Risk-of-Bias Tool for Randomized Trials: Exploratory Feasibility Study. J Med Internet Res. 2026.
- Lin HT, Yeh JT. meta-pipe: AI-assisted, end-to-end meta-analysis pipeline with reproducible tooling [software]. GitHub. commit 5c5c3f0cc2347bd4f991d2071dc68abad5c8a5ab (2026-09-23)
- Nussbaumer-Streit B, Dobrescu A, Sharifan A et al. LLM-Based Classifiers Can Reduce the Proportion of Records to Screen with Minimal Loss of Relevant Studies: A Case Study. J Clin Epidemiol. 2026.
- Pitre T, Zeraatkar D, Granton J et al. A biomedical BERT ensemble (TITAN-SR) outperformed active learning and LLM chatbots for systematic-review title and abstract screening: development and external validation. J Clin Epidemiol. 2026.
- Saran A, Macleod M, Nduku P et al. Safe and responsible use of AI in evidence synthesis: Recommendations from the Evidence Synthesis Infrastructure Collaborative (ESIC) Working Group 3. J Clin Epidemiol. 2026.
- Shankar R, Lim A, Qian X Performance of large language models in data extraction for evidence synthesis: A systematic review. J Biomed Inform. 2026.
- Xie C, Kong W, Pi L et al. Performance of Large Language Models in Automated Medical Literature Screening: A Systematic Review and Meta-Analysis. J Evid Based Med. 2026.
教科書、軟體與教學資源(3 筆)
- Viechtbauer W Conducting Meta-Analyses in R with the metafor Package. Journal of Statistical Software 36(3). 2010.
- Schwarzer G, Carpenter JR, Rucker G Meta-Analysis with R (Springer, book). Springer. 2015.
- Harrer M, Cuijpers P, Furukawa TA, Ebert DD Doing Meta-Analysis with R: A Hands-On Guide. Chapman & Hall/CRC (book). 2021.
AI 時代章節:人工流程基準與其他論證文獻(22 筆)
- Jones AP, et al. High prevalence but low impact of data extraction and reporting errors were found in Cochrane systematic reviews. J Clin Epidemiol. 2005;58:741-2
- Gøtzsche PC, et al. Data extraction errors in meta-analyses that use standardized mean differences. JAMA. 2007;298:430-7
- Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J Am Med Inform Assoc. 2012;19:121-7
- Elliott JH, et al. Living systematic reviews: an emerging opportunity to narrow the evidence-practice gap. PLoS Med. 2014;11:e1001603
- Yavchitz A, et al. A new classification of spin in systematic reviews and meta-analyses was developed and ranked according to the severity. J Clin Epidemiol. 2016;75:56-65
- Borah R, et al. Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry. BMJ Open. 2017;7:e012545
- Mathes T, et al. Frequency of data extraction errors and methods to increase data extraction quality: a methodological review. BMC Med Res Methodol. 2017;17:152
- Thomas J, et al. Living systematic reviews: 2. Combining human and machine effort. J Clin Epidemiol. 2017;91:31-37
- Gates A, Vandermeer B, Hartling L. Technology-assisted risk of bias assessment in systematic reviews: a prospective cross-sectional evaluation of the RobotReviewer machine learning tool. J Clin Epidemiol. 2018;96:54-62
- Gartlehner G, et al. Single-reviewer abstract screening missed 13 percent of relevant studies: a crowd-based, randomized controlled trial. J Clin Epidemiol. 2020;121:20-28
- Minozzi S, et al. The revised Cochrane risk of bias tool for randomized trials (RoB 2) showed low interrater reliability and challenges in its application. J Clin Epidemiol. 2020;126:37-44
- Wang Z, et al. Error rates of human reviewers during abstract screening in systematic reviews. PLoS One. 2020;15:e0227742
- Hu K, et al. Inconsistencies in study eligibility criteria are common between non-Cochrane systematic reviews and their protocols registered in PROSPERO. Res Synth Methods. 2021;12:394-405
- Chen L, Zaharia M, Zou J. How is ChatGPT's behavior changing over time? arXiv:2307.09009. 2023
- Panickssery A, Bowman SR, Feng S. LLM evaluators recognize and favor their own generations. arXiv:2404.13076. 2024
- Kowall B, et al. Marital status and risk of cardiovascular disease — a multi-analyst study in epidemiology. Eur J Epidemiol. 2025
- Oami T, et al. Optimal large language models to screen citations for systematic reviews. Res Synth Methods. 2025
- Taneri PE. Human versus artificial intelligence: comparing Cochrane authors' and ChatGPT's risk of bias assessments. Cochrane Evid Synth Methods. 2025
- Lin HT, Yeh JT. meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis. arXiv:2606.28363. 2026
- Nyrhi L, et al. Large language models for risk-of-bias assessment in randomised clinical trials — a comparative validation study. EBioMedicine. 2026
- Purewal A, et al. Human versus artificial intelligence: evaluating ChatGPT's performance in conducting published systematic reviews with meta-analysis in chronic pain research. Reg Anesth Pain Med. 2026
- Xiong YT, et al. Impact of prompt engineering on large language models for risk of bias assessment: a comparative study. BMJ Evid Based Med. 2026