Benchmark for filter methods for feature selection in high-dimensional classification data

被引:337
|
作者
Bommert, Andrea [1 ]
Sun, Xudong [2 ]
Bischl, Bernd [2 ]
Rahnenfuehrer, Joerg [1 ]
Lang, Michel [1 ]
机构
[1] TU Dortmund Univ, Dept Stat, D-44221 Dortmund, Germany
[2] Ludwig Maximilians Univ Munchen, Dept Stat, Ludwigstr 33, D-80539 Munich, Germany
关键词
Feature selection; Filter methods; High-dimensional data; Benchmark; INFORMATION; ALGORITHMS; RELEVANCE; MODEL;
D O I
10.1016/j.csda.2019.106839
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Feature selection is one of the most fundamental problems in machine learning and has drawn increasing attention due to high-dimensional data sets emerging from different fields like bioinformatics. For feature selection, filter methods play an important role, since they can be combined with any machine learning model and can heavily reduce run time of machine learning algorithms. The aim of the analyses is to review how different filter methods work, to compare their performance with respect to both run time and predictive accuracy, and to provide guidance for applications. Based on 16 high-dimensional classification data sets, 22 filter methods are analyzed with respect to run time and accuracy when combined with a classification method. It is concluded that there is no group of filter methods that always outperforms all other methods, but recommendations on filter methods that perform well on many of the data sets are made. Also, groups of filters that are similar with respect to the order in which they rank the features are found. For the analyses, the R machine learning package mlr is used. It provides a uniform programming API and therefore is a convenient tool to conduct feature selection using filter methods. (C) 2019 The Authors. Published by Elsevier B.V.
引用
收藏
页数:19
相关论文
共 50 条
  • [21] Feature selection for high-dimensional temporal data
    Tsagris, Michail
    Lagani, Vincenzo
    Tsamardinos, Ioannis
    [J]. BMC BIOINFORMATICS, 2018, 19
  • [22] FEATURE SELECTION FOR HIGH-DIMENSIONAL DATA ANALYSIS
    Verleysen, Michel
    [J]. ECTA 2011/FCTA 2011: PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON EVOLUTIONARY COMPUTATION THEORY AND APPLICATIONS AND INTERNATIONAL CONFERENCE ON FUZZY COMPUTATION THEORY AND APPLICATIONS, 2011,
  • [23] Interaction-based feature selection and classification for high-dimensional biological data
    Wang, Haitian
    Lo, Shaw-Hwa
    Zheng, Tian
    Hu, Inchi
    [J]. BIOINFORMATICS, 2012, 28 (21) : 2834 - 2842
  • [24] FACO: A Novel Hybrid Feature Selection Algorithm for High-Dimensional Data Classification
    Popoola, Gideon
    Oyeniran, Kayode
    [J]. SOUTHEASTCON 2024, 2024, : 61 - 68
  • [25] Efficient feature selection for high-dimensional data using two-level filter
    Li, Y
    Wu, ZF
    Liu, JM
    Tang, YY
    [J]. PROCEEDINGS OF THE 2004 INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND CYBERNETICS, VOLS 1-7, 2004, : 1711 - 1716
  • [26] A new improved filter-based feature selection model for high-dimensional data
    Munirathinam, Deepak Raj
    Ranganadhan, Mohanasundaram
    [J]. JOURNAL OF SUPERCOMPUTING, 2020, 76 (08): : 5745 - 5762
  • [27] A new improved filter-based feature selection model for high-dimensional data
    Deepak Raj Munirathinam
    Mohanasundaram Ranganadhan
    [J]. The Journal of Supercomputing, 2020, 76 : 5745 - 5762
  • [28] Hybrid Filter and Genetic Algorithm-Based Feature Selection for Improving Cancer Classification in High-Dimensional Microarray Data
    Ali, Waleed
    Saeed, Faisal
    [J]. PROCESSES, 2023, 11 (02)
  • [30] Efficient feature selection filters for high-dimensional data
    Ferreira, Artur J.
    Figueiredo, Mario A. T.
    [J]. PATTERN RECOGNITION LETTERS, 2012, 33 (13) : 1794 - 1804