Human versus machine in medicine: can scientific literature answer the question?
Tessa S. Cook
- Year
- 2019
- Citations
- 16
Abstract
In The Lancet Digital Health, Xiaoxuan Liu and colleagues1Liu X Faes L Kale A et al.A comparison of deep learning performance against health care profesisonals in detecting diseases from medical imaging: a systematic review and meta-analysis.Lancet Digital Health. 2019; (published online Sept 24)https://doi.org/10.1016/S2589-7500(19)30123-2Summary Full Text Full Text PDF PubMed Scopus (357) Google Scholar present a systematic review and meta-analysis in an attempt to answer the question of whether deep learning is better than human health-care professionals across all imaging domains of medicine. Despite the plethora of headlines proclaiming how the latest artificial intelligence (AI) has outperformed a human physician, the authors found surprisingly few studies that compare the performance of humans and these models. From more than 20 000 unique abstracts, fewer than 100 studies met their eligibility criteria for the systematic review and only 25 met their inclusion criteria for the meta-analysis. These 25 studies compared the performance of deep learning solutions to health-care professionals for 13 different specialty areas, only two of which—breast cancer and dermatological cancers—were represented by more than three studies. The meta-analysis suggests equivalent performance of deep learning algorithms and health-care professionals in the 14 studies that used the same out-of-sample validation dataset to compare their performances, showing a pooled sensitivity of 87·0% (95% CI 83·0–90·2) for deep learning models and 86·4% (79·9–91·0) for health-care professionals, and a pooled specificity of 92·5% (85·1–96·4) for deep learning models and 90·5% (80·6–95·7) for health-care professionals. This work nicely illustrates the challenge of attempting to compare AI with humans for medical applications, and the authors rightly qualify their conclusion with a detailed list of potential confounders and limitations. The eventual sample size representing a broad swath of the domain of medicine underlines the need for a deeper dive into the literature.2Challen R Denny J Pitt M Gompels L Edwards T Tsaneva-Atanasova K Artificial intelligence, bias and clinical safety.BMJ Qual Saf. 2019; 28: 231-237Crossref PubMed Scopus (174) Google Scholar Evaluation of diagnostic accuracy—whether for AI or otherwise—requires a ground truth. In the absence of a perfect ground truth, inherent biases are introduced into a study.3Crawford K Calo R There is a blind spot in AI research.Nature. 2016; 538: 311-313Crossref PubMed Scopus (109) Google Scholar This is particularly problematic when evaluating AI tools to decide if they perform better than humans. As Liu and colleagues point out, there is a wide spectrum of what constitutes expert consensus or ground truth in the literature, yet these datasets with inconsistent, imperfect, or even incorrect labels become training and testing data for AI models. If researchers cannot all agree on what it means to agree, how can we know if model A is better than human B? More importantly, how can an AI model be trained when experts themselves disagree on the correct answer to a question?4Lallas A Argenziano G Artificial intelligence and melanoma diagnosis: ignoring human nature may lead to false predictions.Dermatol Pract Concept. 2018; 8: 249-251Crossref PubMed Google Scholar AI cannot yet replicate the essence of the diagnostic process. In medicine, different datapoints become available at different times during a work-up. One test might be ordered because of the result of another. So, when AI algorithms are trained on a complete corpus of retrospective data that eliminates both the temporal variation and the dependency within the data, can it actually be compared with the human physician who made a series of related decisions to create that comprehensive dataset?5Liang H Tsui BY Ni H Valentim CCS Baxter SL Liu G et al.Evaluation and accurate diagnoses of pediatric diseases using artificial intelligence.Nat Med. 2019; 2
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002