Hepatitis C Virus Detection Model by Using Random Forest, Logistic-Regression and ABC Algorithm

被引:10
|
作者
Li, Tzuu-Hseng S. [1 ]
Chiu, Huan-Jung [1 ]
Kuo, Ping-Huan [2 ]
机构
[1] Natl Cheng Kung Univ, Dept Elect Engn, aiRobots Lab, Tainan 70101, Taiwan
[2] Natl Chung Cheng Univ, Dept Mech Engn, Chiayi 62102, Taiwan
关键词
Liver diseases; Classification tree analysis; Random forests; Data models; Classification algorithms; Artificial bee colony algorithm; Medical diagnostic imaging; Monte Carlo methods; Sampling methods; Random forest; logistic regression; two-stage mixing; ABC algorithm; 10-fold Monte-Carlo cross-validation; synthetic minority oversampling technique; DISEASE DIAGNOSIS; LIVER-DISEASE; CLASSIFICATION; PREDICTION;
D O I
10.1109/ACCESS.2022.3202295
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
This study proposes an automatic classifier for detecting the multiclass probabilities of hepatitis C virus (HCV) incidence based on patients' blood attributes. The purpose of this study is to establish an artificial intelligence-based model that can identify HCV patients and detect the disease in early stage for future treatments. This model can be applied by using clinical data and keeps the performance from imbalanced datasets. The innovation in this article lies in considering the "unbalanced data" existing in medical record-based clinical data. Synthetic minority oversampling technique (SMOTE) algorithm was further employed to derive corresponding solutions. This objective was achieved using a cascade two-stage method combining the random forest (RF) and logistic regression (LR) algorithms. Two models were trained by applying the RF (Model 1) and LR (Model 2) to raw and preprocessed data, respectively. The artificial bee colony (ABC) algorithm was then used to determine the optimal threshold value required for filtering and separation, that is, the optimal combination of both models. The two-stage mixing algorithm combines algorithms of different search dimensions, thus integrating the strengths of those algorithms. The critical threshold value for separating Model 1 and Model 2 was obtained through an optimized search using the ABC algorithm. After conducting 10-fold Monte Carlo cross-validation experiments 50 times (for mean values), data from the recent pandemic were used to verify the proposed method. To evaluate the quantitative results, indicators, such as prediction accuracy, precision, recall, F1-score, and Matthews correlation coefficient, were compared with those of the latest algorithms used in relevant fields. The results indicate that the proposed model, named Cascade RF-LR (with SMOTE), can be used to detect the multiclass probabilities of HCV incidence using the ABC algorithm, thereby improving the effectiveness of relevant treatments.
引用
收藏
页码:91045 / 91058
页数:14
相关论文
共 50 条
  • [21] Classification and Prediction of Heart Disease using Novel Random Forest Algorithm by Comparing Logistic Regression for Obtaining Better Accuracy
    Poojitha, T.
    Mahaveerakannan, R.
    CARDIOMETRY, 2022, (25): : 1538 - 1545
  • [22] Analysis of English Writing Text Features Based on Random Forest and Logistic Regression Classification Algorithm
    Sun, Chuan
    Luo, Bo
    MOBILE INFORMATION SYSTEMS, 2022, 2022
  • [23] A Corrosion Detection Algorithm Via The Random Forest Model
    Liu Tingting
    Kang Kai
    Zhang Fen
    Ni Jialiang
    Wang Tianyun
    17TH INTERNATIONAL CONFERENCE ON OPTICAL COMMUNICATIONS AND NETWORKS (ICOCN2018), 2019, 11048
  • [24] Using Decision Trees, Logistic Regression and Random Forest to Predict Poverty Risk in Thailand
    Meenorngwar, Chai
    2024 12th International Conference on Cyber and IT Service Management, CITSM 2024, 2024,
  • [25] CLASSIFYING HIGH MEDICAL EXPENDITURE PATIENTS USING LOGISTIC REGRESSION AND RANDOM FOREST METHODS
    Menon, J.
    VALUE IN HEALTH, 2021, 24 : S188 - S189
  • [26] Comparison of Accuracy Rate in Prediction of Cardiovascular Disease using Random Forest with Logistic Regression
    Vishnuvardhan, Talluri
    Rama, A.
    CARDIOMETRY, 2022, (25): : 1526 - 1531
  • [27] Random Forest and Logistic Regression algorithms for prediction of groundwater contamination using ammonia concentration
    Ahmed Madani
    Mohammed Hagage
    Salwa F. Elbeih
    Arabian Journal of Geosciences, 2022, 15 (20)
  • [28] VALIDATION OF A LOGISTIC-REGRESSION MODEL FOR GROUP-A BETA-STREP USING ROC CURVE ANALYSIS
    CENTOR, RM
    WIGTON, RS
    CONNOR, JL
    MEDICAL DECISION MAKING, 1983, 3 (03) : 391 - 391
  • [29] Analysis of Spam Detection Using Integration of Logistic Regression and PSO Algorithm
    Ponmalar, A.
    Rajkumar, K.
    Hariharan, U.
    Kalaiselvi, V.K.G.
    Deeba, S.
    Proceedings of the 2021 4th International Conference on Computing and Communications Technologies, ICCCT 2021, 2021, : 396 - 402
  • [30] An Innovative Penalty based Heart Disease Prediction system using Novel Random Forest over Logistic Regression Classifier Algorithm
    Teja, P. Prasanna Sai
    Veeramani, T.
    CARDIOMETRY, 2022, (25): : 1477 - 1482