A Comparative Analysis of Data Mining Techniques on Breast Cancer Diagnosis Data using WEKA Toolbox

被引:0
|
作者
Alshammari, Majdah [1 ]
Mezher, Mohammad [1 ]
机构
[1] Fahad Bin Sultan Univ, Dept Comp Sci, Tabuk, Saudi Arabia
关键词
Data mining; breast cancer; data mining techniques; classification; WEKA toolbox;
D O I
10.14569/IJACSA.2020.0110829
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Breast cancer is considered the second most common cancer in women compared to all other cancers. It is fatal in less than half of all cases and is the main cause of mortality in women. It accounts for 16% of all cancer mortalities worldwide. Early diagnosis of breast cancer increases the chance of recovery. Data mining techniques can be utilized in the early diagnosis of breast cancer. In this paper, an academic experimental breast cancer dataset is used to perform a data mining practical experiment using the Waikato Environment for Knowledge Analysis (WEKA) tool. The WEKA Java application represents a rich resource for conducting performance metrics during the execution of experiments. Pre-processing and feature extraction are used to optimize the data. The classification process used in this study was summarized through thirteen experiments. Additionally, 10 experiments using various different classification algorithms were conducted. The introduced algorithms were: Naive Bayes, Logistic Regression, Lazy IBK (Instance-Bases learning with parameter K), Lazy Kstar, Lazy Locally Weighted Learner, Rules ZeroR, Decision Stump, Decision Trees J48, Random Forest and Random Trees. The process of producing a predictive model was automated with the use of classification accuracy. Further, several experiments on classification of Wisconsin Diagnostic Breast Cancer and Wisconsin Breast Cancer, were conducted to compare the success rates of the different methods. Results conclude that Lazy IBK classifier k-NN can achieve 98% accuracy among other classifiers. The main advantages of the study were the compactness of using 13 different data mining models and 10 different performance measurements, and plotting figures of classifications errors.
引用
收藏
页码:224 / 229
页数:6
相关论文
共 50 条
  • [1] Using Data Mining Techniques to Support Breast Cancer Diagnosis
    Diz, Joana
    Marreiros, Goreti
    Freitas, Alberto
    [J]. NEW CONTRIBUTIONS IN INFORMATION SYSTEMS AND TECHNOLOGIES, VOL 1, PT 1, 2015, 353 : 689 - 700
  • [2] Comparative Analysis of Breast Cancer and Hypothyroid Dataset using Data Mining Classification Techniques
    Verma, Deepika
    Mishra, Nidhi
    [J]. 2017 IEEE INTERNATIONAL CONFERENCE ON POWER, CONTROL, SIGNALS AND INSTRUMENTATION ENGINEERING (ICPCSI), 2017, : 1624 - 1626
  • [3] Comparative analysis of three data mining techniques in diagnosis of lung cancer
    Li, Di
    Li, Zunshui
    Ding, Mingcui
    Ni, Ran
    Wang, Jing
    Qu, Lingbo
    Wang, Wei
    Wu, Yongjun
    [J]. EUROPEAN JOURNAL OF CANCER PREVENTION, 2021, 30 (01) : 15 - 20
  • [4] Analysis of breast cancer using data mining & statistical techniques
    Xiong, XC
    Kim, YO
    Baek, YC
    Rhee, DW
    Kim, SH
    [J]. Sixth International Conference on Software Engineerng, Artificial Intelligence, Networking and Parallel/Distributed Computing and First AICS International Workshop on Self-Assembling Wireless Networks, Proceedings, 2005, : 82 - 87
  • [5] A Survey on Breast Cancer Analysis Using Data Mining Techniques
    Padmapriya, B.
    Velmurugan, T.
    [J]. 2014 IEEE INTERNATIONAL CONFERENCE ON COMPUTATIONAL INTELLIGENCE AND COMPUTING RESEARCH (IEEE ICCIC), 2014, : 1234 - 1237
  • [6] Applying Data Mining Techniques to Improve Breast Cancer Diagnosis
    Joana Diz
    Goreti Marreiros
    Alberto Freitas
    [J]. Journal of Medical Systems, 2016, 40
  • [7] Applying Data Mining Techniques to Improve Breast Cancer Diagnosis
    Diz, Joana
    Marreiros, Goreti
    Freitas, Alberto
    [J]. JOURNAL OF MEDICAL SYSTEMS, 2016, 40 (09)
  • [8] Data mining in bioinformatics using Weka
    Frank, E
    Hall, M
    Trigg, L
    Holmes, G
    Witten, IH
    [J]. BIOINFORMATICS, 2004, 20 (15) : 2479 - 2481
  • [9] Data mining techniques in breast cancer diagnosis at the cellular–molecular level
    Jian Yang
    Dler Hussein Kadir
    [J]. Journal of Cancer Research and Clinical Oncology, 2023, 149 : 12605 - 12620
  • [10] Comparative Analysis of Data Mining Techniques for Business Data
    Jamil, Jastini Mohd
    Shaharanee, Izwan Nizal Mohd
    [J]. INTERNATIONAL CONFERENCE ON QUANTITATIVE SCIENCES AND ITS APPLICATIONS (ICOQSIA 2014), 2014, 1635 : 587 - 593