Multiple imputation in the presence of high-dimensional data

被引:52
|
作者
Zhao, Yize [1 ]
Long, Qi [1 ]
机构
[1] Emory Univ, Dept Biostat & Bioinformat, Atlanta, GA 30322 USA
关键词
Bayesian lasso regression; high-dimensional data; missing data; multiple imputation; regularized regression; FULLY CONDITIONAL SPECIFICATION; MULTIVARIATE IMPUTATION; LASSO ESTIMATORS; REGRESSION; SELECTION;
D O I
10.1177/0962280213511027
中图分类号
R19 [保健组织与事业(卫生事业管理)];
学科分类号
摘要
Missing data are frequently encountered in biomedical, epidemiologic and social research. It is well known that a naive analysis without adequate handling of missing data may lead to bias and/or loss of efficiency. Partly due to its ease of use, multiple imputation has become increasingly popular in practice for handling missing data. However, it is unclear what is the best strategy to conduct multiple imputation in the presence of high-dimensional data. To answer this question, we investigate several approaches of using regularized regression and Bayesian lasso regression to impute missing values in the presence of high-dimensional data. We compare the performance of these methods through numerical studies, in which we also evaluate the impact of the dimension of the data, the size of the true active set for imputation, and the strength of correlation. Our numerical studies show that in the presence of high-dimensional data the standard multiple imputation approach performs poorly and the imputation approach using Bayesian lasso regression achieves, in most cases, better performance than the other imputation methods including the standard imputation approach using the correctly specified imputation model. Our results suggest that Bayesian lasso regression and its extensions are better suited for multiple imputation in the presence of high-dimensional data than the other regression methods.
引用
收藏
页码:2021 / 2035
页数:15
相关论文
共 50 条
  • [1] Multiple Imputation for General Missing Data Patterns in the Presence of High-dimensional Data
    Deng, Yi
    Chang, Changgee
    Ido, Moges Seyoum
    Long, Qi
    [J]. SCIENTIFIC REPORTS, 2016, 6
  • [2] Multiple Imputation for General Missing Data Patterns in the Presence of High-dimensional Data
    Yi Deng
    Changgee Chang
    Moges Seyoum Ido
    Qi Long
    [J]. Scientific Reports, 6
  • [3] Multiple imputation with compatibility for high-dimensional data
    Zahid, Faisal Maqbool
    Faisal, Shahla
    Heumann, Christian
    [J]. PLOS ONE, 2021, 16 (07):
  • [4] Multiple imputation and analysis for high-dimensional incomplete proteomics data
    Yin, Xiaoyan
    Levy, Daniel
    Willinger, Christine
    Adourian, Aram
    Larson, Martin G.
    [J]. STATISTICS IN MEDICINE, 2016, 35 (08) : 1315 - 1326
  • [5] Missing Data Imputation with High-Dimensional Data
    Brini, Alberto
    van den Heuvel, Edwin R.
    [J]. AMERICAN STATISTICIAN, 2024, 78 (02): : 240 - 252
  • [6] Variable selection techniques after multiple imputation in high-dimensional data
    Faisal Maqbool Zahid
    Shahla Faisal
    Christian Heumann
    [J]. Statistical Methods & Applications, 2020, 29 : 553 - 580
  • [7] Multiple imputation for high-dimensional mixed incomplete continuous and binary data
    He, Ren
    Belin, Thomas
    [J]. STATISTICS IN MEDICINE, 2014, 33 (13) : 2251 - 2262
  • [8] Variable selection techniques after multiple imputation in high-dimensional data
    Zahid, Faisal Maqbool
    Faisal, Shahla
    Heumann, Christian
    [J]. STATISTICAL METHODS AND APPLICATIONS, 2020, 29 (03): : 553 - 580
  • [9] Bootstrap-multiple-imputation; high-dimensional model validation with missing data
    Chang, Billy
    Demetrashvili, Nino
    Kowgier, Matthew
    [J]. CANADIAN JOURNAL OF STATISTICS-REVUE CANADIENNE DE STATISTIQUE, 2011, 39 (02): : 202 - 204
  • [10] Imputation of rounded zeros for high-dimensional compositional data
    Templ, Matthias
    Hron, Karel
    Filzmoser, Peter
    Gardlo, Alzbeta
    [J]. CHEMOMETRICS AND INTELLIGENT LABORATORY SYSTEMS, 2016, 155 : 183 - 190