The Effect of Different Dimensionality Reduction Techniques on Machine Learning Overfitting Problem

被引:0
|
作者
Salam, Mustafa Abdul [1 ]
Azar, Ahmad Taher [2 ,3 ]
Elgendy, Mustafa Samy [4 ]
Fouad, Khaled Mohamed [5 ]
机构
[1] Benha Univ, Fac Comp & Artificial Intelligence, Artificial Intelligence Dept, Banha, Egypt
[2] Benha Univ, Fac Comp & Artificial Intelligence, Banha, Egypt
[3] Prince Sultan Univ, Coll Comp & Informat Sci, Riyadh, Saudi Arabia
[4] Benha Univ, Sci Comp Dept, Fac Comp & Artificial Intelligence, Banha, Egypt
[5] Benha Univ, Informat Syst Dept, Fac Comp & Artificial Intelligence, Banha, Egypt
关键词
Dimensionality reduction; feature subset selection; rough set; overfitting; underfitting; machine learning; PRINCIPAL COMPONENT ANALYSIS; ALGORITHM;
D O I
暂无
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
In most conditions, it is a problematic mission for a machine-learning model with a data record, which has various attributes, to be trained. There is always a proportional relationship between the increase of model features and the arrival to the overfitting of the susceptible model. That observation occurred since not all the characteristics are always important. For example, some features could only cause the data to be noisier. Dimensionality reduction techniques are used to overcome this matter. This paper presents a detailed comparative study of nine dimensionality reduction methods. These methods are missing-values ratio, low variance filter, highcorrelation filter, random forest, principal component analysis, linear discriminant analysis, backward feature elimination, forward feature construction, and rough set theory. The effects of used methods on both training and testing performance were compared with two different datasets and applied to three different models. These models are, Artificial Neural Network (ANN), Support Vector Machine (SVM) and Random Forest classifier (RFC). The results proved that the RFC model was able to achieve the dimensionality reduction via limiting the overfitting crisis. The introduced RFC model showed a general progress in both accuracy and efficiency against compared approaches. The results revealed that dimensionality reduction could minimize the overfitting process while holding the performance so near to or better than the original one.
引用
收藏
页码:641 / 655
页数:15
相关论文
共 50 条
  • [41] Dimensionality Reduction, Modelling, and Optimization of Multivariate Problems Based on Machine Learning
    Alswaitti, Mohammed
    Siddique, Kamran
    Jiang, Shulei
    Alomoush, Waleed
    Alrosan, Ayat
    [J]. SYMMETRY-BASEL, 2022, 14 (07):
  • [42] A Dimensionality Reduction Approach for Machine Learning Based IoT Botnet Detection
    Susanto
    Stiawan, Deris
    Arifin, M. Agus Syamsul
    Rejito, Juli
    Idris, Mohd. Yazid
    Budiarto, Rahmat
    [J]. 2021 8TH INTERNATIONAL CONFERENCE ON ELECTRICAL ENGINEERING, COMPUTERSCIENCE AND INFORMATICS (EECSI) 2021, 2021, : 26 - 30
  • [43] Machine learning methods for nonlinear dimensionality reduction of the thermospheric density field
    Nateghi, Vahid
    Manzi, Matteo
    [J]. ADVANCES IN SPACE RESEARCH, 2023, 72 (10) : 4106 - 4114
  • [44] Fast and Reliable DDoS Detection using Dimensionality Reduction and Machine Learning
    Ashi, Zein
    Aburashed, Laila
    Al-Fawa'reh, Mohammad
    Qasaimeh, Malek
    [J]. INTERNATIONAL CONFERENCE FOR INTERNET TECHNOLOGY AND SECURED TRANSACTIONS (ICITST-2020), 2020, : 13 - 22
  • [45] A novel dimensionality reduction approach by integrating dynamics theory and machine learning
    Chen, Xiyuan
    Wang, Qiubao
    [J]. MATHEMATICS AND COMPUTERS IN SIMULATION, 2024, 218 : 98 - 111
  • [46] Spatial Correlation Preserving EEG Dimensionality Reduction Using Machine Learning
    Gebre-Amlak, Haymanot
    Nguyen, Hoang
    Lowe, Jesse
    Nabulsi, Ala-Addin
    Chu, Narisa Nan
    [J]. PROCEEDINGS 2018 IEEE INTERNATIONAL CONFERENCE ON BIOINFORMATICS AND BIOMEDICINE (BIBM), 2018, : 2583 - 2589
  • [47] A Novel Framework for Sentiment Analysis: Dimensionality Reduction for Machine Learning (DRML)
    Dhamayanthi, N.
    Lavanya, B.
    [J]. INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2024, 15 (06) : 777 - 794
  • [48] Wind turbine fault detection: a semi-supervised learning approach with two different dimensionality reduction techniques
    de Sá F.P.G.
    de Coutinho R.C.
    Ogasawara E.
    Brandão D.
    Toso R.F.
    [J]. International Journal of Innovative Computing and Applications, 2023, 14 (1-2) : 67 - 77
  • [49] Improving asphalt mix design by predicting alligator cracking and longitudinal cracking based on machine learning and dimensionality reduction techniques
    Liu, Jian
    Liu, Fangyu
    Gong, Hongren
    Fanijo, Ebenezer O.
    Wang, Linbing
    [J]. CONSTRUCTION AND BUILDING MATERIALS, 2022, 354
  • [50] A multi-stage method to predict carbon dioxide emissions using dimensionality reduction, clustering, and machine learning techniques
    Mardani, Abbas
    Liao, Huchang
    Nilashi, Mehrbakhsh
    Alrasheedi, Melfi
    Cavallaro, Fausto
    [J]. JOURNAL OF CLEANER PRODUCTION, 2020, 275