XAI-reduct: accuracy preservation despite dimensionality reduction for heart disease classification using explainable AI

被引：9

作者：

Das, Surajit ^{[1
,2
]}

Sultana, Mahamuda ^{[3
]}

Bhattacharya, Suman ^{[2
]}

Sengupta, Diganta ^{[4
]}

De, Debashis ^{[2
]}

机构：

[1] Meghnad Saha Inst Technol, Dept Informat Technol, Kolkata 700150, India

[2] Maulana Abul Kalam Azad Univ Technol, Dept Comp Sci & Engn, Nadia 741249, West Bengal, India

[3] Guru Nanak Inst Technol, Dept Comp Sci & Engn, Kolkata 700114, India

[4] Meghnad Saha Inst Technol, Dept Comp Sci & Engn, Kolkata 700150, India

来源：

JOURNAL OF SUPERCOMPUTING | 2023年 / 79卷 / 16期

关键词：

Explainable machine learning; Heart disease classification; SHAP; LIME; SHAPASH; DALEX; PDP; XAI; Dimensionality reduction;

D O I：

10.1007/s11227-023-05356-3

中图分类号：

TP3 [计算技术、计算机技术];

学科分类号：

0812 ;

摘要：

Machine learning (ML) has been used for classification of heart diseases for almost a decade, although understanding of the internal working of the black boxes, i.e., non-interpretable models, remain a demanding problem. Another major challenge in such ML models is the curse of dimensionality leading to resource intensive classification using the comprehensive set of feature vector (CFV). This study focuses on dimensionality reduction using explainable artificial intelligence, without negotiating on accuracy for heart disease classification. Four explainable ML models, using SHAP, were used for classification which reflected the feature contributions (FC) and feature weights (FW) for each feature in the CFV for generating the final results. FC and FW were taken into account in generating the reduced dimensional feature subset (FS). The findings of the study are as follows: (a) XGBoost classifies heart diseases best with explanations, with an increase in 2% in model accuracy over existing best proposals, (b) explainable classification using FS exhibits better accuracy than most of the literary proposals, and (c) with the increase in explainability, accuracy can be preserved using XGBoost classifier for classifying heart diseases, and (d) the top four features responsible for diagnosis of heart disease have been exhibited which have common occurrences in all the explanations reflected by the five explainable techniques used on XGBoost classifier based on feature contributions. To the best of our knowledge, this is first attempt to explain XGBoost classification for diagnosis of heart diseases using five explainable techniques.

引用

页码：18167 / 18197

页数：31