A Simple Information Criterion for Variable Selection in High-Dimensional Regression

被引:0
|
作者
Pluntz, Matthieu [1 ]
Dalmasso, Cyril [2 ]
Tubert-Bitter, Pascale [1 ]
Ahmed, Ismail [1 ]
机构
[1] Univ Paris Sud, Univ Paris Saclay, High Dimens Biostat Drug Safety & Genom, UVSQ,Inserm,CESP, Villejuif, France
[2] Univ Evry Val Essonne, Lab Math & Modelisat Evry LaMME, Evry, France
基金
中国国家自然科学基金;
关键词
FWER control; high-dimensional regression; information criterion; LASSO; pharmacovigilance; variable selection; MODEL SELECTION; REGULARIZATION; LIKELIHOOD; RISK;
D O I
10.1002/sim.10275
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
High-dimensional regression problems, for example with genomic or drug exposure data, typically involve automated selection of a sparse set of regressors. Penalized regression methods like the LASSO can deliver a family of candidate sparse models. To select one, there are criteria balancing log-likelihood and model size, the most common being AIC and BIC. These two methods do not take into account the implicit multiple testing performed when selecting variables in a high-dimensional regression, which makes them too liberal. We propose the extended AIC (EAIC), a new information criterion for sparse model selection in high-dimensional regressions. It allows for asymptotic FWER control when the candidate regressors are independent. It is based on a simple formula involving model log-likelihood, model size, the total number of candidate regressors, and the FWER target. In a simulation study over a wide range of linear and logistic regression settings, we assessed the variable selection performance of the EAIC and of other information criteria (including some that also use the number of candidate regressors: mBIC, mAIC, and EBIC) in conjunction with the LASSO. Our method controls the FWER in nearly all settings, in contrast to the AIC and BIC, which produce many false positives. We also illustrate it for the automated signal detection of adverse drug reactions on the French pharmacovigilance spontaneous reporting database.
引用
收藏
页数:12
相关论文
共 50 条
  • [41] Variable selection for high-dimensional incomplete data
    Liang, Lixing
    Zhuang, Yipeng
    Yu, Philip L. H.
    COMPUTATIONAL STATISTICS & DATA ANALYSIS, 2024, 192
  • [42] Variable selection in high-dimensional partially linear additive models for composite quantile regression
    Guo, Jie
    Tang, Manlai
    Tian, Maozai
    Zhu, Kai
    COMPUTATIONAL STATISTICS & DATA ANALYSIS, 2013, 65 : 56 - 67
  • [43] SCAD-penalized quantile regression for high-dimensional data analysis and variable selection
    Amin, Muhammad
    Song, Lixin
    Thorlie, Milton Abdul
    Wang, Xiaoguang
    STATISTICA NEERLANDICA, 2015, 69 (03) : 212 - 235
  • [44] Sparse Bayesian variable selection in high-dimensional logistic regression models with correlated priors
    Ma, Zhuanzhuan
    Han, Zifei
    Ghosh, Souparno
    Wu, Liucang
    Wang, Min
    STATISTICAL ANALYSIS AND DATA MINING, 2024, 17 (01)
  • [45] SPATIAL BAYESIAN VARIABLE SELECTION AND GROUPING FOR HIGH-DIMENSIONAL SCALAR-ON-IMAGE REGRESSION
    Li, Fan
    Zhang, Tingting
    Wang, Quanli
    Gonzalez, Marlen Z.
    Maresh, Erin L.
    Coan, James A.
    ANNALS OF APPLIED STATISTICS, 2015, 9 (02): : 687 - 713
  • [46] Variable selection in the single-index quantile regression model with high-dimensional covariates
    Kuruwita, C. N.
    COMMUNICATIONS IN STATISTICS-SIMULATION AND COMPUTATION, 2023, 52 (03) : 1120 - 1132
  • [47] Transfer learning for sparse variable selection in high-dimensional regression from quadratic measurement
    Shang, Qingxu
    Li, Jie
    Song, Yunquan
    KNOWLEDGE-BASED SYSTEMS, 2024, 300
  • [48] High-dimensional graphs and variable selection with the Lasso
    Meinshausen, Nicolai
    Buehlmann, Peter
    ANNALS OF STATISTICS, 2006, 34 (03): : 1436 - 1462
  • [49] High-Dimensional Variable Selection for Survival Data
    Ishwaran, Hemant
    Kogalur, Udaya B.
    Gorodeski, Eiran Z.
    Minn, Andy J.
    Lauer, Michael S.
    JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2010, 105 (489) : 205 - 217
  • [50] FAITHFUL VARIABLE SCREENING FOR HIGH-DIMENSIONAL CONVEX REGRESSION
    Xu, Min
    Chen, Minhua
    Lafferty, John
    ANNALS OF STATISTICS, 2016, 44 (06): : 2624 - 2660