Multiclass Classification on High Dimension and Low Sample Size Data Using Genetic Programming

被引:3
|
作者
Wei, Tingyang [1 ]
Liu, Wei-Li [1 ]
Zhong, Jinghui [1 ]
Gong, Yue-Jiao [1 ]
机构
[1] South China Univ Technol, Sch Comp Sci & Engn, Guangzhou 510006, Peoples R China
关键词
Machine learning; Feature extraction; Gene expression; Programming; Genetic programming; Sociology; Statistics; gene expression programming; high dimension; classification; low sample size; ensemble learning; FEATURE-SELECTION; RULES;
D O I
10.1109/TETC.2020.3034495
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Multiclass classification is one of the most fundamental tasks in data mining. However, traditional data mining methods rely on the model assumption, they generally can suffer from the overfitting problem on high dimension and low sample size (HDLSS) data. Trying to address multiclass classification problems on HDLSS data from another perspective, we utilize Genetic Programming (GP), an intrinsic evolutionary classification algorithm that can implement feature construction automatically without model assumption. This article develops an ensemble-based genetic programming classification framework, the Sigmoid-based Ensemble Gene Expression Programming (SE-GEP). To relieve the problem of output conflict in GP-based multiclass classifiers, the proposed method employs a flexible probability representation with continuous relaxation to better integrate the output of all the binary classifiers, an effective data division strategy to further enhance the ensemble performance, and a novel sampling strategy to refine the existing GP-based binary classifier. The experiment results indicate that SE-GEP can attain better classification accuracy compared to other GP methods. Moreover, the comparison with other representative machine learning methods indicates that SE-GEP is a competitive method for multiclass classification in HDLSS data.
引用
收藏
页码:704 / 718
页数:15
相关论文
共 50 条
  • [1] Classification for high-dimension low-sample size data
    Shen, Liran
    Er, Meng Joo
    Yin, Qingbo
    [J]. PATTERN RECOGNITION, 2022, 130
  • [2] Classification for high-dimension low-sample size data
    Shen, Liran
    Er, Meng Joo
    Yin, Qingbo
    [J]. PATTERN RECOGNITION, 2022, 130
  • [3] Some considerations of classification for high dimension low-sample size data
    Zhang, Lingsong
    Lin, Xihong
    [J]. STATISTICAL METHODS IN MEDICAL RESEARCH, 2013, 22 (05) : 537 - 550
  • [4] Multiclass object classification using genetic programming
    Zhang, MJ
    Smart, W
    [J]. APPLICATIONS OF EVOLUTIONARY COMPUTING, 2004, 3005 : 369 - 378
  • [5] On some transformations of high dimension, low sample size data for nearest neighbor classification
    Dutta, Subhajit
    Ghosh, Anil K.
    [J]. MACHINE LEARNING, 2016, 102 (01) : 57 - 83
  • [6] On some transformations of high dimension, low sample size data for nearest neighbor classification
    Subhajit Dutta
    Anil K. Ghosh
    [J]. Machine Learning, 2016, 102 : 57 - 83
  • [7] On Perfect Clustering of High Dimension, Low Sample Size Data
    Sarkar, Soham
    Ghosh, Anil K.
    [J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2020, 42 (09) : 2257 - 2272
  • [8] Geometric representation of high dimension, low sample size data
    Hall, P
    Marron, JS
    Neeman, A
    [J]. JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY, 2005, 67 : 427 - 444
  • [9] CLUSTERING HIGH DIMENSION, LOW SAMPLE SIZE DATA USING THE MAXIMAL DATA PILING DISTANCE
    Ahn, Jeongyoun
    Lee, Myung Hee
    Yoon, Young Joo
    [J]. STATISTICA SINICA, 2012, 22 (02) : 443 - 464
  • [10] Genetic Programming Based ECOC for Multiclass Microarray Data Classification
    Wang JiaJun
    Liu KunHong
    Sun MengXin
    Hong QingQi
    [J]. 2017 10TH INTERNATIONAL SYMPOSIUM ON COMPUTATIONAL INTELLIGENCE AND DESIGN (ISCID), VOL. 1, 2017, : 280 - 283