Longitudinal Nonresponse Prediction with Time Series Machine Learning

被引：0

作者：

Collins, John ^{[1
]}

Kern, Christoph ^{[2
]}

机构：

[1] Univ Mannheim, Mannheimer Zent Europa Sozialforschung, Mannheim, Germany

[2] Ludwig Maximilian Univ Munich, Social Data Sci & AI Lab, Munich, Germany

来源：

JOURNAL OF SURVEY STATISTICS AND METHODOLOGY | 2024年

关键词：

Catch22; Machine learning; Nonresponse; Panel attrition; Recurrent neural network; Time series classification; PANEL ATTRITION; HOUSEHOLD; BIAS;

D O I：

10.1093/jssam/smae037

中图分类号：

O1 [数学]; C [社会科学总论];

学科分类号：

03 ; 0303 ; 0701 ; 070101 ;

摘要：

Panel surveys are an important tool for social science researchers, but nonresponse in any panel wave can significantly reduce data quality. Panel managers then attempt to identify participants who may be at risk of not participating using predictive models to target interventions before data collection through adaptive designs. Previous research has shown that these predictions can be improved by accounting for a sample member's behavior in past waves. These past behaviors are often operationalized through rolling average variables that aggregate information over the past two, three, or all waves, such as each participant's nonresponse rate. However, it is possible that this approach is too simple. In this paper, we evaluate models that account for more nuanced temporal dependency, namely recurrent neural networks (RNNs) and feature-, interval-, and kernel-based time series classification techniques. We compare these novel techniques' performances to more traditional logistic regression and tree-based models in predicting future panel survey nonresponse. We apply these algorithms to predict nonresponse in the GESIS Panel, a large-scale, probability-based German longitudinal study, for surveys conducted between 2013 and 2021. Our findings show that RNNs perform similar to tree-based approaches, but the RNNs do not require the analyst to create rolling average variables. More complex feature-, interval-, and kernel-based techniques are not more effective at classifying future respondents and nonrespondents than RNNs or traditional logistic regression or tree-based methods. We find that predicting nonresponse of newly recruited participants is a more difficult task, and basic RNN models and penalized logistic regression performed best in this situation. We conclude that RNNs may be better at classifying future response propensity than traditional logistic regression and tree-based approaches when the association between time-varying characteristics and survey participation is complex but did not do so in the current analysis when a traditional rolling averages approach yielded comparable results.

引用

页数：32

共 50 条

[41] Online sequential extreme learning machine with kernels for nonstationary time series prediction
Wang, Xinying
Han, Min
[J]. NEUROCOMPUTING, 2014, 145 : 90 - 97
[42] Time-series failure prediction on small datasets using machine learning
Maior, Caio B. S.
Silva, Thaylon G.
[J]. IEEE LATIN AMERICA TRANSACTIONS, 2024, 22 (05) : 362 - 371
[43] Multivariate Time Series Prediction based on Multiple Kernel Extreme Learning Machine
Wang, Xinying
Han, Min
[J]. PROCEEDINGS OF THE 2014 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2014, : 198 - 201
[44] Short Term Prediction of Continuous Time Series Based on Extreme Learning Machine
Wang, Hongbo
Song, Peng
Wang, Chengyao
Tu, Xuyan
[J]. PROCEEDINGS OF ELM-2016, 2018, 9 : 113 - 127
[45] Human pol II promoter prediction: time series descriptors and machine learning
Gangal, R
Sharma, P
[J]. NUCLEIC ACIDS RESEARCH, 2005, 33 (04) : 1332 - 1336
[46] Stock Price Prediction Using Time Series, Econometric, Machine Learning, and Deep Learning Models
Chatterjee, Ananda
Bhowmick, Hrisav
Sen, Jaydip
[J]. 2021 IEEE Mysore Sub Section International Conference, MysuruCon 2021, 2021, : 289 - 296
[47] Stock price prediction using time series, econometric, machine learning, and deep learning models
Chatterjee, Ananda
Bhowmick, Hrisav
Sen, Jaydip
[J]. arXiv, 2021,
[48] A propensity score adjustment method for longitudinal time series models under nonignorable nonresponse
Liu, Zhan
Yau, Chun Yip
[J]. STATISTICAL PAPERS, 2022, 63 (01) : 317 - 342
[49] A propensity score adjustment method for longitudinal time series models under nonignorable nonresponse
Zhan Liu
Chun Yip Yau
[J]. Statistical Papers, 2022, 63 : 317 - 342
[50] Learning and prediction of relational time series
Tan, Terence K.
Darken, Christian J.
[J]. COMPUTATIONAL AND MATHEMATICAL ORGANIZATION THEORY, 2015, 21 (02) : 210 - 241

← 1 2 3 4 5 →