A Combined Feature Screening Approach of Random Forest and Filter-based Methods for Ultra-high Dimensional Data

被引:10
|
作者
Zhou, Lifeng [1 ]
Wang, Hong [2 ]
机构
[1] Changsha Univ, Sch Econ & Management, Changsha, Peoples R China
[2] Cent South Univ, Sch Math & Stat, Changsha, Peoples R China
关键词
Feature screening; filter-based method; ultra-high dimensional data; variable selection; random forest; RF-DCSIS; BREAST-CANCER; FEATURE-SELECTION; CELL-PROLIFERATION; VARIABLE SELECTION; EXPRESSION; PROGNOSIS; SCUBE2; MIGRATION;
D O I
10.2174/1574893617666220221120618
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background: Various feature (variable) screening approaches have been proposed in the past decade to mitigate the impact of ultra-high dimensionality in classification and regression problems, including filter based methods such as sure independence screening, and wrapper based methods such as random forest. However, the former type of methods rely heavily on strong modelling assumptions while the latter ones requires an adequate sample size to make the data speak for themselves. These requirements can seldom be met in biochemical studies in cases where we have only access to ultra-high dimensional data with a complex structure and a small number of observations. Objective: In this research, we want to investigate the possibility of combining both filter based screening methods and random forest based screening methods in the regression context. Methods: We have combined four state-of-art filter approaches, namely, sure independence screening (SIS), robust rank correlation based screening (RRCS), high dimensional ordinary least squares projection (HOLP) and a model free sure independence screening procedure based on the distance correlation (DCSIS) from the statistical community with a random forest based Boruta screening method from the machine learning community for regression problems. Results: Among all the combined methods, RF-DCSIS performs better than the other methods in terms of screening accuracy and prediction capability on the simulated scenarios and real benchmark datasets. Conclusion: By empirical study from both extensive simulation and real data, we have shown that both filter based screening and random forest based screening have their pros and cons, while a combination of both may lead to a better feature screening result and prediction capability.
引用
收藏
页码:344 / 357
页数:14
相关论文
共 50 条
  • [1] Generalized Jaccard feature screening for ultra-high dimensional survival data
    Liu, Renqing
    Deng, Guangming
    He, Hanji
    [J]. AIMS MATHEMATICS, 2024, 9 (10): : 27607 - 27626
  • [2] Grouped feature screening for ultra-high dimensional data for the classification model
    He, Hanji
    Deng, Guangming
    [J]. JOURNAL OF STATISTICAL COMPUTATION AND SIMULATION, 2022, 92 (05) : 974 - 997
  • [3] Feature Selection in High Dimensional Data by a Filter-Based Genetic Algorithm
    De Stefano, Claudio
    Fontanella, Francesco
    di Freca, Alessandra Scotto
    [J]. APPLICATIONS OF EVOLUTIONARY COMPUTATION, EVOAPPLICATIONS 2017, PT I, 2017, 10199 : 506 - 521
  • [4] Sequential Feature Screening for Generalized Linear Models with Sparse Ultra-High Dimensional Data
    Junying Zhang
    Hang Wang
    Riquan Zhang
    Jiajia Zhang
    [J]. Journal of Systems Science and Complexity, 2020, 33 : 510 - 526
  • [5] Model-free feature screening for ultra-high dimensional competing risks data
    Chen, Xiaolin
    Zhang, Yahui
    Liu, Yi
    Chen, Xiaojing
    [J]. STATISTICS & PROBABILITY LETTERS, 2020, 164
  • [6] Sequential Feature Screening for Generalized Linear Models with Sparse Ultra-High Dimensional Data
    ZHANG Junying
    WANG Hang
    ZHANG Riquan
    ZHANG Jiajia
    [J]. Journal of Systems Science & Complexity, 2020, 33 (02) : 510 - 526
  • [7] Sequential Feature Screening for Generalized Linear Models with Sparse Ultra-High Dimensional Data
    Zhang, Junying
    Wang, Hang
    Zhang, Riquan
    Zhang, Jiajia
    [J]. JOURNAL OF SYSTEMS SCIENCE & COMPLEXITY, 2020, 33 (02) : 510 - 526
  • [8] Adjusted feature screening for ultra-high dimensional missing response
    Zou, Liying
    Liu, Yi
    Zhang, Zhonghu
    [J]. JOURNAL OF STATISTICAL COMPUTATION AND SIMULATION, 2024, 94 (03) : 460 - 483
  • [9] A new improved filter-based feature selection model for high-dimensional data
    Munirathinam, Deepak Raj
    Ranganadhan, Mohanasundaram
    [J]. JOURNAL OF SUPERCOMPUTING, 2020, 76 (08): : 5745 - 5762
  • [10] A new improved filter-based feature selection model for high-dimensional data
    Deepak Raj Munirathinam
    Mohanasundaram Ranganadhan
    [J]. The Journal of Supercomputing, 2020, 76 : 5745 - 5762