Active Learning for Handling Missing Data

被引：0

作者：

Tharwat, Alaa ^{[1
]}

Schenck, Wolfram ^{[1
]}

机构：

[1] Univ Appl Sci & Arts, Hsch Bielefeld, Ctr Appl Data Sci CfADS, D-33619 Bielefeld, Germany

来源：

IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS | 2024年

关键词：

Uncertainty; Labeling; Data models; Training data; Costs; Predictive models; Search problems; Active learning (AL); imputation uncertainty; missing data; multiple imputation;

D O I：

10.1109/TNNLS.2024.3352279

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Recently, the massive growth of IoT devices and Internet data, which are widely used in many applications, including industry and healthcare, has dramatically increased the amount of free unlabeled data collected. However, this unlabeled data is useless if we want to learn supervised machine learning models. The expensive and time-consuming cost of labeling makes the problem even more challenging. Here, the active learning (AL) technique provides a solution by labeling small but highly informative and representative data, which guarantees a high degree of generalizability over space and improves classification performance with data we have never seen before. The task is more difficult when the active learner has no predefined knowledge, such as initial training data, and when the obtained data is incomplete (i.e., contains missing values). In previous studies, the missing data should first be imputed. Then, the active learner selects from the available unlabeled data, regardless of whether the points were originally observed or imputed. However, selecting inaccurate imputed data points would negatively affect the active learner and prevent it from selecting informative and/or representative points, thus reducing the overall classification performance of the prediction models. This motivated us to introduce a novel query selection strategy that accounts for imputation uncertainty when querying new points. For this purpose, we first introduce a novel multiple imputation method that considers feature importance in selecting the most promising feature groups for missing values estimation. This multiple imputation method provides the ability to quantify the imputation uncertainty of each imputed data point. Furthermore, in each of the two phases of the proposed active learner (exploration and exploitation), imputation uncertainty is taken into account to reduce the probability of selecting points with high imputation uncertainty. We tested the effectiveness of the proposed active learner on different binary and multiclass datasets with different missing rates.

引用

页码：1 / 15

页数：15

共 50 条

[41] Handling Missing Data in Randomized Experiments with Noncompliance
Booil Jo
Elizabeth M. Ginexi
Nicholas S. Ialongo
[J]. Prevention Science, 2010, 11 : 384 - 396
[42] Methods for Handling Missing Secondary Respondent Data
Young, Rebekah
Johnson, David
[J]. JOURNAL OF MARRIAGE AND FAMILY, 2013, 75 (01) : 221 - 234
[43] Comparison of Methods for Handling Missing Covariate Data
Åsa M. Johansson
Mats O. Karlsson
[J]. The AAPS Journal, 2013, 15 : 1232 - 1241
[44] Modern concepts in the handling and reporting of missing data
Lee, Katherine
Carpenter, James
Little, Roderick
Nguyen, Cattram
Cornish, Rosie
[J]. INTERNATIONAL JOURNAL OF EPIDEMIOLOGY, 2021, 50
[45] Taxonomy of Missing Data along with their handling Methods
Tripathi, Ashok Kumar
Rathee, Geetanjali
Saini, Hemraj
[J]. 2019 FIFTH INTERNATIONAL CONFERENCE ON IMAGE INFORMATION PROCESSING (ICIIP 2019), 2019, : 463 - 468
[46] Comparison of Methods for Handling Missing Covariate Data
Johansson, Asa M.
Karlsson, Mats O.
[J]. AAPS JOURNAL, 2013, 15 (04): : 1232 - 1241
[47] Handling Missing Data in Randomized Experiments with Noncompliance
Jo, Booil
Ginexi, Elizabeth M.
Ialongo, Nicholas S.
[J]. PREVENTION SCIENCE, 2010, 11 (04) : 384 - 396
[48] Learning with Missing Data
Escobar, Carlos A.
Arinez, Jorge
Macias, Daniela
Morales-Menendez, Ruben
[J]. 2020 IEEE INTERNATIONAL CONFERENCE ON BIG DATA (BIG DATA), 2020, : 5037 - 5045
[49] A framework for handling missing accelerometer outcome data in trials
Tackney, Mia S.
Cook, Derek G.
Stahl, Daniel
Ismail, Khalida
Williamson, Elizabeth
Carpenter, James
[J]. TRIALS, 2021, 22 (01)
[50] Handling missing covariate data in clinical studies in haematology
Bonneville, Edouard F.
Schetelig, Johannes
Putter, Hein
de Wreede, Liesbeth C.
[J]. BEST PRACTICE & RESEARCH CLINICAL HAEMATOLOGY, 2023, 36 (02)

← 1 2 3 4 5 →