Multiple imputation in the presence of high-dimensional data

被引：52

作者：

Zhao, Yize ^{[1
]}

Long, Qi ^{[1
]}

机构：

[1] Emory Univ, Dept Biostat & Bioinformat, Atlanta, GA 30322 USA

来源：

STATISTICAL METHODS IN MEDICAL RESEARCH | 2016年 / 25卷 / 05期

关键词：

Bayesian lasso regression; high-dimensional data; missing data; multiple imputation; regularized regression; FULLY CONDITIONAL SPECIFICATION; MULTIVARIATE IMPUTATION; LASSO ESTIMATORS; REGRESSION; SELECTION;

D O I：

10.1177/0962280213511027

中图分类号：

R19 [保健组织与事业（卫生事业管理）];

学科分类号：

摘要：

Missing data are frequently encountered in biomedical, epidemiologic and social research. It is well known that a naive analysis without adequate handling of missing data may lead to bias and/or loss of efficiency. Partly due to its ease of use, multiple imputation has become increasingly popular in practice for handling missing data. However, it is unclear what is the best strategy to conduct multiple imputation in the presence of high-dimensional data. To answer this question, we investigate several approaches of using regularized regression and Bayesian lasso regression to impute missing values in the presence of high-dimensional data. We compare the performance of these methods through numerical studies, in which we also evaluate the impact of the dimension of the data, the size of the true active set for imputation, and the strength of correlation. Our numerical studies show that in the presence of high-dimensional data the standard multiple imputation approach performs poorly and the imputation approach using Bayesian lasso regression achieves, in most cases, better performance than the other imputation methods including the standard imputation approach using the correctly specified imputation model. Our results suggest that Bayesian lasso regression and its extensions are better suited for multiple imputation in the presence of high-dimensional data than the other regression methods.

引用

页码：2021 / 2035

页数：15

共 50 条

[41] High-dimensional data visualization
Lin Tang
Nature Methods, 2020, 17 : 129 - 129
[42] High-dimensional Data Cubes
John, Sachin Basil
Koch, Christoph
PROCEEDINGS OF THE VLDB ENDOWMENT, 2022, 15 (13): : 3828 - 3840
[43] Modeling High-Dimensional Data
Vempala, Santosh S.
COMMUNICATIONS OF THE ACM, 2012, 55 (02) : 112 - 112
[44] Learning high-dimensional data
Verleysen, M
LIMITATIONS AND FUTURE TRENDS IN NEURAL COMPUTATION, 2003, 186 : 141 - 162
[45] A telescope for high-dimensional data
Shneiderman, B
COMPUTING IN SCIENCE & ENGINEERING, 2006, 8 (02) : 48 - 53
[46] Clustering High-Dimensional Data
Masulli, Francesco
Rovetta, Stefano
CLUSTERING HIGH-DIMENSIONAL DATA, CHDD 2012, 2015, 7627 : 1 - 13
[47] High-Dimensional Data in Genomics
Amaratunga, Dhammika
Cabrera, Javier
BIOPHARMACEUTICAL APPLIED STATISTICS SYMPOSIUM, VOL 3: PHARMACEUTICAL APPLICATIONS, 2018, : 65 - 73
[48] PLOTS OF HIGH-DIMENSIONAL DATA
ANDREWS, DF
BIOMETRICS, 1972, 28 (01) : 125 - &
[49] A shortcut to high-dimensional data
Nature Methods, 2018, 15 (1) : 15 - 15
[50] Optimal multiple change-point detection for high-dimensional data
Pilliat, Emmanuel
Carpentier, Alexandra
Verzelen, Nicolas
ELECTRONIC JOURNAL OF STATISTICS, 2023, 17 (01): : 1240 - 1315

← 1 2 3 4 5 →