Knockoff boosted tree for model-free variable selection

被引：7

作者：

Jiang, Tao ^{[1
]}

Li, Yuanyuan ^{[2
]}

Motsinger-Reif, Alison A. ^{[2
]}

机构：

[1] North Carolina State Univ, Bioinformat Res Ctr, Dept Stat, Raleigh, NC 27695 USA

[2] NIEHS, Biostat & Computat Biol Branch, Durham, NC 27709 USA

来源：

BIOINFORMATICS | 2021年 / 37卷 / 07期

关键词：

MOLECULAR CLASSIFICATION; EXPRESSION; REGRESSION; CANCER; DISCOVERY; GENES;

D O I：

10.1093/bioinformatics/btaa770

中图分类号：

Q5 [生物化学];

学科分类号：

071010 ; 081704 ;

摘要：

Motivation: The recently proposed knockoff filter is a general framework for controlling the false discovery rate (FDR) when performing variable selection. This powerful new approach generates a 'knockoff' of each variable tested for exact FDR control. Imitation variables that mimic the correlation structure found within the original variables serve as negative controls for statistical inference. Current applications of knockoff methods use linear regression models and conduct variable selection only for variables existing in model functions. Here, we extend the use of knockoffs for machine learning with boosted trees, which are successful and widely used in problems where no prior knowledge of model function is required. However, currently available importance scores in tree models are insufficient for variable selection with FDR control. Results: We propose a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. We extend the current knockoff method to model-free variable selection through the use of tree-based models. Additionally, we propose and evaluate two new sampling methods for generating knockoffs, namely the sparse covariance and principal component knockoff methods. We test and compare these methods with the original knockoff method regarding their ability to control type I errors and power. In simulation tests, we compare the properties and performance of importance test statistics of tree models. The results include different combinations of knockoffs and importance test statistics. We consider scenarios that include main-effect, interaction, exponential and second-order models while assuming the true model structures are unknown. We apply our algorithm for tumor purity estimation and tumor classification using Cancer Genome Atlas (TCGA) gene expression data. Our results show improved discrimination between difficult-to-discriminate cancer types.

引用

页码：976 / 983

页数：8

共 50 条

[1] Model-free variable selection
Li, LX
Cook, RD
Nachtsheim, CJ
[J]. JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY, 2005, 67 : 285 - 299
[2] Model-free variable selection for conditional mean in regression
Dong, Yuexiao
Yu, Zhou
Zhu, Liping
[J]. COMPUTATIONAL STATISTICS & DATA ANALYSIS, 2020, 152
[3] MODEL-FREE SELECTION OF ORDERED AND CONTINUOUS VARIABLE COMBINATIONS
TARTER, ME
HILBERMANN, M
KAMM, B
[J]. AMERICAN JOURNAL OF EPIDEMIOLOGY, 1974, 100 (06) : 529 - 529
[4] Model-Free Feature Screening and FDR Control With Knockoff Features
Liu, Wanjun
Ke, Yuan
Liu, Jingyuan
Li, Runze
[J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2022, 117 (537) : 428 - 443
[5] On dual model-free variable selection with two groups of variables
Alothman, Ahmad
Dong, Yuexiao
Artemiou, Andreas
[J]. JOURNAL OF MULTIVARIATE ANALYSIS, 2018, 167 : 366 - 377
[6] Shrinkage inverse regression estimation for model-free variable selection
Bondell, Howard D.
Li, Lexin
[J]. JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY, 2009, 71 : 287 - 299
[7] Trace Pursuit: A General Framework for Model-Free Variable Selection
Yu, Zhou
Dong, Yuexiao
Zhu, Li-Xing
[J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2016, 111 (514) : 813 - 821
[8] Multiple Loci Mapping via Model-free Variable Selection
Sun, Wei
Li, Lexin
[J]. BIOMETRICS, 2012, 68 (01) : 12 - 22
[9] Model-Free Variable Selection With Matrix-Valued Predictors
Li, Zeda
Dong, Yuexiao
[J]. JOURNAL OF COMPUTATIONAL AND GRAPHICAL STATISTICS, 2021, 30 (01) : 171 - 181
[10] Model-free Variable Selection in Reproducing Kernel Hilbert Space
Yang, Lei
Lv, Shaogao
Wang, Junhui
[J]. JOURNAL OF MACHINE LEARNING RESEARCH, 2016, 17

← 1 2 3 4 5 →