A new hybrid ensemble machine-learning model for severity risk assessment and post-COVID prediction system

被引：15

作者：

Shakhovska, Natalya ^{[1
]}

Yakovyna, Vitaliy ^{[1
,2
]}

Chopyak, Valentyna ^{[3
]}

机构：

[1] Lviv Polytech Natl Univ, Dept Artificial Intelligence, UA-79013 Lvov, Ukraine

[2] Univ Warmia & Mazury, Fac Math & Comp Sci, PL-10719 Olsztyn, Poland

[3] Danylo Halytskyi Lviv Natl Univ, Dept Clin Immunol & Allergol, UA-79010 Lvov, Ukraine

来源：

MATHEMATICAL BIOSCIENCES AND ENGINEERING | 2022年 / 19卷 / 06期

基金：

新加坡国家研究基金会;

关键词：

COVID-19; severity prediction; machine learning; ensemble classification; biomarkers; FEATURE-SELECTION;

D O I：

10.3934/mbe.2022285

中图分类号：

Q [生物科学];

学科分类号：

07 ; 0710 ; 09 ;

摘要：

Starting from December 2019, the COVID-19 pandemic has globally strained medical resources and caused significant mortality. It is commonly recognized that the severity of SARS-CoV-2 disease depends on both the comorbidity and the state of the patient's immune system, which is reflected in several biomarkers. The development of early diagnosis and disease severity prediction methods can reduce the burden on the health care system and increase the effectiveness of treatment and rehabilitation of patients with severe cases. This study aims to develop and validate an ensemble machine-learning model based on clinical and immunological features for severity risk assessment and post-COVID rehabilitation duration for SARS-CoV-2 patients. The dataset consisting of 35 features and 122 instances was collected from Lviv regional rehabilitation center. The dataset contains age, gender, weight, height, BMI, CAT, 6-minute walking test, pulse, external respiration function, oxygen saturation, and 15 immunological markers used to predict the relationship between disease duration and biomarkers using the machine learning approach. The predictions are assessed through an area under the receiver-operating curve, classification accuracy, precision, recall, and F1 score performance metrics. A new hybrid ensemble feature selection model for a post-COVID prediction system is proposed as an automatic feature cut-off rank identifier. A three-layer high accuracy stacking ensemble classification model for intelligent analysis of short medical datasets is presented. Together with weak predictors, the associative rules allowed improving the classification quality. The proposed ensemble allows using a random forest model as an aggregator for weak repressors' results generalization. The performance of the three-layer stacking ensemble classification model (AUC 0.978; CA 0.920; F1 score 0.921; precision 0.924; recall 0.920) was higher than five machine learning models, viz. tree algorithm with forward pruning; Naive Bayes classifier; support vector machine with RBF kernel; logistic regression, and a calibrated learner with sigmoid function and decision threshold optimization. Aging-related biomarkers, viz. CD3+, CD4+, CD8+, CD22+ were examined to predict post-COVID rehabilitation duration. The best accuracy was reached in the case of the support vector machine with the linear kernel (MAPE = 0.0787) and random forest classifier (RMSE = 1.822). The proposed three -layer stacking ensemble classification model predicted SARS-CoV-2 disease severity based on the cytokines and physiological biomarkers. The results point out that changes in studied biomarkers associated with the severity of the disease can be used to monitor the severity and forecast the rehabilitation duration.

引用

页码：6102 / 6123

页数：22

共 50 条

[41] Commentary: Machine learning and the brave new world of risk model assessment
Kurlansky, Paul
JOURNAL OF THORACIC AND CARDIOVASCULAR SURGERY, 2023, 165 (04): : 1445 - 1446
[42] New hybrid data mining model for prediction of Salmonella presence in agricultural waters based on ensemble feature selection and machine learning algorithms
Buyrukoglu, Selim
JOURNAL OF FOOD SAFETY, 2021, 41 (04)
[43] Explainable XGBoost-SHAP Machine-Learning Model for Prediction of Ground Motion Duration in New Zealand
Somala, Surendra Nadh
Chanda, Sarit
Alhamaydeh, Mohammad
Mangalathu, Sujith
NATURAL HAZARDS REVIEW, 2024, 25 (02)
[44] Toward explainable flood risk prediction: Integrating a novel hybrid machine learning model
Wang, Yongyang
Zhang, Pan
Xie, Yulei
Chen, Lei
Li, Yu
SUSTAINABLE CITIES AND SOCIETY, 2025, 120
[45] Developing the breast cancer risk prediction system using hybrid machine learning algorithms
Afrash, Mohammad R.
Bayani, Azadeh
Shanbehzadeh, Mostafa
Bahadori, Mohammadkarim
Kazemi-Arpanahi, Hadi
JOURNAL OF EDUCATION AND HEALTH PROMOTION, 2022, 11 (01) : 272
[46] Water distillation tower: Experimental investigation, economic assessment, and performance prediction using optimized machine-learning model
Elsheikh, Ammar H.
El-Said, Emad M. S.
Abd Elaziz, Mohamed
Fujii, Manabu
El-Tahan, Hamed R.
JOURNAL OF CLEANER PRODUCTION, 2023, 388
[47] Explainable Machine Learning for Early Assessment of COVID-19 Risk Prediction in Emergency Departments
Casiraghi, Elena
Malchiodi, Dario
Trucco, Gabriella
Frasca, Marco
Cappelletti, Luca
Fontana, Tommaso
Esposito, Alessandro Andrea
Avola, Emanuele
Jachetti, Alessandro
Reese, Justin
Rizzi, Alessandro
Robinson, Peter N.
Valentini, Giorgio
IEEE ACCESS, 2020, 8 (08): : 196299 - 196325
[48] Development and validation of a hybrid deep learning-machine learning approach for severity assessment of COVID-19 and other pneumonias
Park, Doohyun
Jang, Ryoungwoo
Chung, Myung Jin
An, Hyun Joon
Bak, Seongwon
Choi, Euijoon
Hwang, Dosik
SCIENTIFIC REPORTS, 2023, 13 (01)
[49] A novel, machine-learning model for prediction of short-term ASCVD risk over 90 and 365 days
Gazit, Tomer
Mann, Hanan
Gaber, Shiri
Adamenko, Pavel
Pariente, Granit
Volsky, Liron
Dolev, Amir
Lyson, Helena
Zimlichman, Eyal
Pandit, Jay A.
Paz, Edo
FRONTIERS IN DIGITAL HEALTH, 2024, 6
[50] Stacking Ensemble-Based Intelligent Machine Learning Model for Predicting Post-COVID-19 Complications
Gupta, Aditya
Jain, Vibha
Singh, Amritpal
NEW GENERATION COMPUTING, 2022, 40 (04) : 987 - 1007

← 1 2 3 4 5 →