Reproducible Machine Learning Methods for Lung Cancer Detection Using Computed Tomography Images: Algorithm Development and Validation

被引：31

作者：

Yu, Kun-Hsing ^{[1
,2
,3
]}

Lee, Tsung-Lu Michael ^{[4
]}

Yen, Ming-Hsuan ^{[5
,6
,7
]}

Kou, S. C. ^{[2
]}

Rosen, Bruce ^{[8
,9
]}

Chiang, Jung-Hsien ^{[7
]}

Kohane, Isaac S. ^{[1
,9
]}

机构：

[1] Harvard Med Sch, Dept Biomed Informat, Boston, MA 02115 USA

[2] Harvard Univ, Dept Stat, Cambridge, MA 02138 USA

[3] Brigham & Womens Hosp, Dept Pathol, 75 Francis St, Boston, MA 02115 USA

[4] Kun Shan Univ, Dept Informat Engn, Tainan, Taiwan

[5] Natl Cheng Kung Univ, Grad Program Multimedia Syst & Intelligent Comp, Tainan, Taiwan

[6] Acad Sinica, Tainan, Taiwan

[7] Natl Cheng Kung Univ, Dept Comp Sci & Informat Engn, 1 Univ Rd, Tainan, Taiwan

[8] Massachusetts Gen Hosp, Dept Radiol, Athinoula A Martinos Ctr Biomed Imaging, Boston, MA USA

[9] Harvard Massachusetts Inst Technol, Div Hlth Sci & Technol, Boston, MA USA

来源：

JOURNAL OF MEDICAL INTERNET RESEARCH | 2020年 / 22卷 / 08期

基金：

美国国家科学基金会; 美国国家卫生研究院;

关键词：

computed tomography; spiral; lung cancer; machine learning; early detection of cancer; reproducibility of results; PULMONARY NODULES; AIDED DIAGNOSIS; CT; RADIOLOGISTS;

D O I：

10.2196/16709

中图分类号：

R19 [保健组织与事业（卫生事业管理）];

学科分类号：

摘要：

Background: Chest computed tomography (CT) is crucial for the detection of lung cancer, and many automated CT evaluation methods have been proposed. Due to the divergent software dependencies of the reported approaches, the developed methods are rarely compared or reproduced. Objective: The goal of the research was to generate reproducible machine learning modules for lung cancer detection and compare the approaches and performances of the award-winning algorithms developed in the Kaggle Data Science Bowl. Methods: We obtained the source codes of all award-winning solutions of the Kaggle Data Science Bowl Challenge, where participants developed automated CT evaluation methods to detect lung cancer (training set n=1397, public test set n=198, final test set n=506). The performance of the algorithms was evaluated by the log-loss function, and the Spearman correlation coefficient of the performance in the public and final test sets was computed. Results: Most solutions implemented distinct image preprocessing, segmentation, and classification modules. Variants of U-Net, VGGNet, and residual net were commonly used in nodule segmentation, and transfer learning was used in most of the classification algorithms. Substantial performance variations in the public and final test sets were observed (Spearman correlation coefficient = .39 among the top 10 teams). To ensure the reproducibility of results, we generated a Docker container for each of the top solutions. Conclusions: We compared the award-winning algorithms for lung cancer detection and generated reproducible Docker images for the top solutions. Although convolutional neural networks achieved decent accuracy, there is plenty of room for improvement regarding model generalizability.

引用

页数：11

共 50 条

[1] Ensemble methods for computed tomography scan images to improve lung cancer detection and classification
Quasar, Syeda Reeha
Sharma, Rishika
Mittal, Aayushi
Sharma, Moolchand
Agarwal, Deevyankar
de La Torre Diez, Isabel
[J]. MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 83 (17) : 52867 - 52897
[2] Ensemble methods for computed tomography scan images to improve lung cancer detection and classification
Syeda Reeha Quasar
Rishika Sharma
Aayushi Mittal
Moolchand Sharma
Deevyankar Agarwal
Isabel de La Torre Díez
[J]. Multimedia Tools and Applications, 2024, 83 : 52867 - 52897
[3] Machine Learning with Data Science-Enabled Lung Cancer Diagnosis and Classification Using Computed Tomography Images
Kiran, S. Vishwa
Kaur, Inderjeet
Thangaraj, K.
Saveetha, V.
Grace, R. Kingsy
Arulkumar, N.
[J]. INTERNATIONAL JOURNAL OF IMAGE AND GRAPHICS, 2023, 23 (03)
[4] Development of a Convolutional Neural Network for detection of Lung Cancer based on Computed Tomography Images
Narvaez, Gabriela
Tirado-Espin, Andres
Cadena-Morejon, Carolina
Villalba-Meneses, Fernando
Cruz-Varela, Jonathan
Villavicencio Gordon, Gabriela
Guevara, Cesar
Alvarado-Cando, Omar
Almeida-Galarraga, Diego
[J]. 2023 FOURTH INTERNATIONAL CONFERENCE ON INFORMATION SYSTEMS AND SOFTWARE TECHNOLOGIES, ICI2ST 2023, 2023, : 24 - 31
[5] Development of Machine Learning Based Algorithm for Prediction of Invasiveness of Early-Lung Adenocarcinoma by Using Chest Computed Tomography
Lee, Juyoung
Park, Seong Yong
Kim, Jin Sung
[J]. MEDICAL PHYSICS, 2021, 48 (06)
[6] Lung Segmentation Using Support Vector Machine in Computed Tomography Images
Molina, Valentin
Vera, Miguel
Robles, Horderlin V.
Bejarano, Edwar
Davila, Hermann
[J]. 2014 PAN AMERICAN HEALTH CARE EXCHANGES (PAHCE), 2014,
[7] A Review of the Detection of Pulmonary Embolism from Computed Tomography Images Using Deep Learning Methods
Das, Manas Pratim
Rohini, V.
[J]. AMBIENT INTELLIGENCE IN HEALTH CARE, ICAIHC 2022, 2023, 317 : 349 - 360
[8] Development and Validation of a Deep Learning Algorithm to Differentiate Colon Carcinoma From Acute Diverticulitis in Computed Tomography Images
Ziegelmayer, Sebastian
Reischl, Stefan
Havrda, Hannah
Gawlitza, Joshua
Graf, Markus
Lenhart, Nicolas
Nehls, Nadja
Lemke, Tristan
Wilhelm, Dirk
Lohoefer, Fabian
Burian, Egon
Neumann, Philipp-Alexander
Makowski, Marcus
Braren, Rickmer
[J]. JAMA NETWORK OPEN, 2023, 6 (01) : E2253370
[9] Using Deep Learning for Classification of Lung Nodules on Computed Tomography Images
Song, QingZeng
Zhao, Lei
Luo, XingKe
Dou, XueChen
[J]. JOURNAL OF HEALTHCARE ENGINEERING, 2017, 2017
[10] Lung Nodule Classification on Computed Tomography Images Using Deep Learning
Amrita Naik
Damodar Reddy Edla
[J]. Wireless Personal Communications, 2021, 116 : 655 - 690

← 1 2 3 4 5 →