Active Learning for Handling Missing Data

被引:0
|
作者
Tharwat, Alaa [1 ]
Schenck, Wolfram [1 ]
机构
[1] Univ Appl Sci & Arts, Hsch Bielefeld, Ctr Appl Data Sci CfADS, D-33619 Bielefeld, Germany
关键词
Uncertainty; Labeling; Data models; Training data; Costs; Predictive models; Search problems; Active learning (AL); imputation uncertainty; missing data; multiple imputation;
D O I
10.1109/TNNLS.2024.3352279
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Recently, the massive growth of IoT devices and Internet data, which are widely used in many applications, including industry and healthcare, has dramatically increased the amount of free unlabeled data collected. However, this unlabeled data is useless if we want to learn supervised machine learning models. The expensive and time-consuming cost of labeling makes the problem even more challenging. Here, the active learning (AL) technique provides a solution by labeling small but highly informative and representative data, which guarantees a high degree of generalizability over space and improves classification performance with data we have never seen before. The task is more difficult when the active learner has no predefined knowledge, such as initial training data, and when the obtained data is incomplete (i.e., contains missing values). In previous studies, the missing data should first be imputed. Then, the active learner selects from the available unlabeled data, regardless of whether the points were originally observed or imputed. However, selecting inaccurate imputed data points would negatively affect the active learner and prevent it from selecting informative and/or representative points, thus reducing the overall classification performance of the prediction models. This motivated us to introduce a novel query selection strategy that accounts for imputation uncertainty when querying new points. For this purpose, we first introduce a novel multiple imputation method that considers feature importance in selecting the most promising feature groups for missing values estimation. This multiple imputation method provides the ability to quantify the imputation uncertainty of each imputed data point. Furthermore, in each of the two phases of the proposed active learner (exploration and exploitation), imputation uncertainty is taken into account to reduce the probability of selecting points with high imputation uncertainty. We tested the effectiveness of the proposed active learner on different binary and multiclass datasets with different missing rates.
引用
收藏
页码:1 / 15
页数:15
相关论文
共 50 条
  • [21] Comparing Methods for Handling Missing Data
    Roda, Celina
    Nicolis, Ioannis
    Momas, Isabelle
    Guihenneuc-Jouyaux, Chantal
    [J]. EPIDEMIOLOGY, 2013, 24 (03) : 469 - 471
  • [22] Handling missing data from heteroskedastic and nonstationary data
    Nelwamondo, Fulufhelo V.
    Marwala, Tshilidzi
    [J]. ADVANCES IN NEURAL NETWORKS - ISNN 2007, PT 1, PROCEEDINGS, 2007, 4491 : 1293 - +
  • [23] A study of handling missing data methods for big data
    Ezzine, Imane
    Benhlima, Laila
    [J]. 2018 IEEE 5TH INTERNATIONAL CONGRESS ON INFORMATION SCIENCE AND TECHNOLOGY (IEEE CIST'18), 2018, : 498 - 501
  • [24] Handling Missing Data in the Modeling of Intensive Longitudinal Data
    Ji, Linying
    Chow, Sy-Miin
    Schermerhom, Alice C.
    Jacobson, Nicholas C.
    Cummings, E. Mark
    [J]. STRUCTURAL EQUATION MODELING-A MULTIDISCIPLINARY JOURNAL, 2018, 25 (05) : 715 - 736
  • [25] Missing values handling for machine learning portfolios
    Chen, Andrew Y.
    McCoy, Jack
    [J]. JOURNAL OF FINANCIAL ECONOMICS, 2024, 155
  • [26] Missing Data Handling using Machine Learning for Human Activity Recognition on Mobile Device
    Prabowo, Okyza M.
    Mutijarsa, Kusprasapta
    Supangkat, Suhono Harso
    [J]. 2016 INTERNATIONAL CONFERENCE ON ICT FOR SMART SOCIETY (ICISS), 2016, : 59 - 62
  • [27] Handling high-dimensional data with missing values by modern machine learning techniques
    Chen, Sixia
    Xu, Chao
    [J]. JOURNAL OF APPLIED STATISTICS, 2023, 50 (03) : 786 - 804
  • [28] A Comparative Study on Missing Data Handling Using Machine Learning for Human Activity Recognition
    Hossain, Tahera
    Inoue, Sozo
    [J]. 2019 JOINT 8TH INTERNATIONAL CONFERENCE ON INFORMATICS, ELECTRONICS & VISION (ICIEV) AND 2019 3RD INTERNATIONAL CONFERENCE ON IMAGING, VISION & PATTERN RECOGNITION (ICIVPR) WITH INTERNATIONAL CONFERENCE ON ACTIVITY AND BEHAVIOR COMPUTING (ABC), 2019, : 124 - 129
  • [29] Handling missing data in diaries of alcohol consumption
    Longford, NT
    Ely, M
    Hardy, R
    Wadsworth, MEJ
    [J]. JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES A-STATISTICS IN SOCIETY, 2000, 163 : 381 - 402
  • [30] Handling missing data in clinical trials: An overview
    Myers, WR
    [J]. DRUG INFORMATION JOURNAL, 2000, 34 (02): : 525 - 533