Semiparametric methods for response-selective and missing data problems in regression

被引:162
|
作者
Lawless, JF
Kalbfleisch, JD
Wild, CJ
机构
[1] Univ Waterloo, Fac Math, Dept Stat & Actuarial Sci, Waterloo, ON N2L 3G1, Canada
[2] Univ Auckland, Auckland 1, New Zealand
关键词
biased sampling; estimated likelihood; estimation; incomplete data; pseudolikelihood;
D O I
10.1111/1467-9868.00185
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
Suppose that data are generated according to the model f(y\x; theta) g(x), where y is a response and x are covariates. We derive and compare semiparametric likelihood and pseudo-likelihood methods for estimating a for situations in which units generated are not fully observed and in which it is impossible or undesirable to model the covariate distribution. The probability that a unit is fully observed may depend on y, and there may be a subset of covariates which is observed only for a subsample of individuals. Our key assumptions are that the probability that a unit has missing data depends only on which of a finite number of strata that (y, x) belongs to and that the stratum membership is observed for every unit. Applications include case-control studies in epidemiology, field reliability studies and broad classes of missing data and measurement error problems. Our results make fully efficient estimation of theta feasible, and they generalize and provide insight into a variety of methods that have been proposed for specific problems.
引用
收藏
页码:413 / 438
页数:26
相关论文
共 50 条