Mining Frequent Itemsets in Correlated Uncertain Databases

被引:0
|
作者
Yong-Xin Tong
Lei Chen
Jieying She
机构
[1] Beihang University,State Key Laboratory of Software Development Environment, School of Computer Science and Engineering
[2] The Hong Kong University of Science and Technology,Department of Computer Science and Engineering
关键词
correlation; uncertain data; probabilistic frequent itemset;
D O I
暂无
中图分类号
学科分类号
摘要
Recently, with the growing popularity of Internet of Things (IoT) and pervasive computing, a large amount of uncertain data, e.g., RFID data, sensor data, real-time video data, has been collected. As one of the most fundamental issues of uncertain data mining, uncertain frequent pattern mining has attracted much attention in database and data mining communities. Although there have been some solutions for uncertain frequent pattern mining, most of them assume that the data is independent, which is not true in most real-world scenarios. Therefore, current methods that are based on the independent assumption may generate inaccurate results for correlated uncertain data. In this paper, we focus on the problem of mining frequent itemsets over correlated uncertain data, where correlation can exist in any pair of uncertain data objects (transactions). We propose a novel probabilistic model, called Correlated Frequent Probability model (CFP model) to represent the probability distribution of support in a given correlated uncertain dataset. Based on the distribution of support derived from the CFP model, we observe that some probabilistic frequent itemsets are only frequent in several transactions with high positive correlation. In particular, the itemsets, which are global probabilistic frequent, have more significance in eliminating the influence of the existing noise and correlation in data. In order to reduce redundant frequent itemsets, we further propose a new type of patterns, called global probabilistic frequent itemsets, to identify itemsets that are always frequent in each group of transactions if the whole correlated uncertain database is divided into disjoint groups based on their correlation. To speed up the mining process, we also design a dynamic programming solution, as well as two pruning and bounding techniques. Extensive experiments on both real and synthetic datasets verify the effectiveness and efficiency of the proposed model and algorithms.
引用
收藏
页码:696 / 712
页数:16
相关论文
共 50 条
  • [1] Mining Frequent Itemsets in Correlated Uncertain Databases
    Tong, Yong-Xin
    Chen, Lei
    She, Jieying
    [J]. JOURNAL OF COMPUTER SCIENCE AND TECHNOLOGY, 2015, 30 (04) : 696 - 712
  • [2] Mining Frequent Itemsets over Uncertain Databases
    Tong, Yongxin
    Chen, Lei
    Cheng, Yurong
    Yu, Philip S.
    [J]. PROCEEDINGS OF THE VLDB ENDOWMENT, 2012, 5 (11): : 1650 - 1661
  • [3] Efficient Mining of Weighted Frequent Itemsets in Uncertain Databases
    Lin, Jerry Chun-Wei
    Gan, Wensheng
    Fournier-Viger, Philippe
    Hong, Tzung-Pei
    [J]. MACHINE LEARNING AND DATA MINING IN PATTERN RECOGNITION (MLDM 2016), 2016, 9729 : 236 - 250
  • [4] Mining Probabilistic Frequent Closed Itemsets in Uncertain Databases
    Tang, Peiyi
    Peterson, Erich A.
    [J]. PROCEEDINGS OF THE 49TH ANNUAL ASSOCIATION FOR COMPUTING MACHINERY SOUTHEAST CONFERENCE (ACMSE '11), 2011, : 86 - 91
  • [5] On Efficient Mining of Frequent Itemsets from Big Uncertain Databases
    Shah, Ahsan
    Halim, Zahid
    [J]. JOURNAL OF GRID COMPUTING, 2019, 17 (04) : 831 - 850
  • [6] On Efficient Mining of Frequent Itemsets from Big Uncertain Databases
    Ahsan Shah
    Zahid Halim
    [J]. Journal of Grid Computing, 2019, 17 : 831 - 850
  • [7] Comprehensive mining of frequent itemsets for a combination of certain and uncertain databases
    Wazir S.
    Beg M.M.S.
    Ahmad T.
    [J]. International Journal of Information Technology, 2020, 12 (4) : 1205 - 1216
  • [8] Mining Weighted Frequent Itemsets without Candidate Generation in Uncertain Databases
    Lin, Jerry Chun-Wei
    Gan, Wensheng
    Fournier-Viger, Philippe
    Hong, Tzung-Pei
    Chao, Han-Chieh
    [J]. INTERNATIONAL JOURNAL OF INFORMATION TECHNOLOGY & DECISION MAKING, 2017, 16 (06) : 1549 - 1579
  • [9] Mining Frequent Itemsets from Multidimensional Databases
    Bay Vo
    Bac Le
    Nguyen, Thang N.
    [J]. INTELLIGENT INFORMATION AND DATABASE SYSTEMS, ACIIDS 2011, PT I, 2011, 6591 : 177 - 186
  • [10] Mining frequent itemsets in distributed and dynamic databases
    Otey, ME
    Wang, C
    Parthasarathy, S
    Veloso, A
    Meira, W
    [J]. THIRD IEEE INTERNATIONAL CONFERENCE ON DATA MINING, PROCEEDINGS, 2003, : 617 - 620