Multiclass Classification on High Dimension and Low Sample Size Data Using Genetic Programming

被引:3
|
作者
Wei, Tingyang [1 ]
Liu, Wei-Li [1 ]
Zhong, Jinghui [1 ]
Gong, Yue-Jiao [1 ]
机构
[1] South China Univ Technol, Sch Comp Sci & Engn, Guangzhou 510006, Peoples R China
关键词
Machine learning; Feature extraction; Gene expression; Programming; Genetic programming; Sociology; Statistics; gene expression programming; high dimension; classification; low sample size; ensemble learning; FEATURE-SELECTION; RULES;
D O I
10.1109/TETC.2020.3034495
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Multiclass classification is one of the most fundamental tasks in data mining. However, traditional data mining methods rely on the model assumption, they generally can suffer from the overfitting problem on high dimension and low sample size (HDLSS) data. Trying to address multiclass classification problems on HDLSS data from another perspective, we utilize Genetic Programming (GP), an intrinsic evolutionary classification algorithm that can implement feature construction automatically without model assumption. This article develops an ensemble-based genetic programming classification framework, the Sigmoid-based Ensemble Gene Expression Programming (SE-GEP). To relieve the problem of output conflict in GP-based multiclass classifiers, the proposed method employs a flexible probability representation with continuous relaxation to better integrate the output of all the binary classifiers, an effective data division strategy to further enhance the ensemble performance, and a novel sampling strategy to refine the existing GP-based binary classifier. The experiment results indicate that SE-GEP can attain better classification accuracy compared to other GP methods. Moreover, the comparison with other representative machine learning methods indicates that SE-GEP is a competitive method for multiclass classification in HDLSS data.
引用
收藏
页码:704 / 718
页数:15
相关论文
共 50 条
  • [21] Using genetic programming for multiclass classification by simultaneously solving component binary classification problems
    Smart, W
    Zhang, MJ
    [J]. GENETIC PROGRAMMING, PROCEEDINGS, 2005, 3447 : 227 - 239
  • [22] A dimension reduction technique applied to regression on high dimension, low sample size neurophysiological data sets
    Adrielle C. Santana
    Adriano V. Barbosa
    Hani C. Yehia
    Rafael Laboissière
    [J]. BMC Neuroscience, 22
  • [23] A dimension reduction technique applied to regression on high dimension, low sample size neurophysiological data sets
    Santana, Adrielle C.
    Barbosa, Adriano V.
    Yehia, Hani C.
    Laboissiere, Rafael
    [J]. BMC NEUROSCIENCE, 2021, 22 (01)
  • [24] A survey of high dimension low sample size asymptotics
    Aoshima, Makoto
    Shen, Dan
    Shen, Haipeng
    Yata, Kazuyoshi
    Zhou, Yi-Hui
    Marron, J. S.
    [J]. AUSTRALIAN & NEW ZEALAND JOURNAL OF STATISTICS, 2018, 60 (01) : 4 - 19
  • [25] A Multiclass Classifier Using Genetic Programming
    Chaudhari, Narendra S.
    Purohit, Anuradha
    Tiwari, Aruna
    [J]. 2008 10TH INTERNATIONAL CONFERENCE ON CONTROL AUTOMATION ROBOTICS & VISION: ICARV 2008, VOLS 1-4, 2008, : 1884 - +
  • [26] Distance-based outlier detection for high dimension, low sample size data
    Ahn, Jeongyoun
    Lee, Myung Hee
    Lee, Jung Ae
    [J]. JOURNAL OF APPLIED STATISTICS, 2019, 46 (01) : 13 - 29
  • [27] Biobjective gradient descent for feature selection on high dimension, low sample size data
    Issa, Tina
    Angel, Eric
    Zehraoui, Farida
    [J]. PLOS ONE, 2024, 19 (07):
  • [28] Discriminating Tensor Spectral Clustering for High-Dimension-Low-Sample-Size Data
    Hu, Yu
    Qi, Fei
    Cheung, Yiu-Ming
    Cai, Hongmin
    [J]. IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2024,
  • [29] Statistical Significance of Clustering for High-Dimension, Low-Sample Size Data
    Liu, Yufeng
    Hayes, David Neil
    Nobel, Andrew
    Marron, J. S.
    [J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2008, 103 (483) : 1281 - 1293
  • [30] Improving Genetic Programming Classification For Binary And Multiclass Datasets
    Al-Madi, Nailah
    Ludwig, Simone A.
    [J]. 2013 IEEE SYMPOSIUM ON COMPUTATIONAL INTELLIGENCE AND DATA MINING (CIDM), 2013, : 166 - 173