Decomposition of variation of mixed variables by a latent mixed Gaussian copula model

被引:1
|
作者
Liu, Yutong [1 ]
Darville, Toni [2 ]
Zheng, Xiaojing [1 ,2 ]
Li, Quefeng [1 ]
机构
[1] Univ N Carolina, Dept Biostat, Chapel Hill, NC 27515 USA
[2] Univ N Carolina, Dept Pediat, Chapel Hill, NC 27515 USA
基金
美国国家卫生研究院;
关键词
high-dimensional matrix estimation; Kendall's tau; latent Gaussian copula model; variation decomposition; GRAPHICAL MODEL; INFECTION; NUMBER; CELLS; JOINT; AXIS; SETS;
D O I
10.1111/biom.13660
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
Many biomedical studies collect data of mixed types of variables from multiple groups of subjects. Some of these studies aim to find the group-specific and the common variation among all these variables. Even though similar problems have been studied by some previous works, their methods mainly rely on the Pearson correlation, which cannot handle mixed data. To address this issue, we propose a latent mixed Gaussian copula (LMGC) model that can quantify the correlations among binary, ordinal, continuous, and truncated variables in a unified framework. We also provide a tool to decompose the variation into the group-specific and the common variation over multiple groups via solving a regularized M-estimation problem. We conduct extensive simulation studies to show the advantage of our proposed method over the Pearson correlation-based methods. We also demonstrate that by jointly solving the M-estimation problem over multiple groups, our method is better than decomposing the variation group by group. We also apply our method to a Chlamydia trachamatis genital tract infection study to demonstrate how it can be used to discover informative biomarkers that differentiate patients.
引用
收藏
页码:1187 / 1200
页数:14
相关论文
共 50 条