Distributed Bayesian Matrix Decomposition for Big Data Mining and Clustering

被引:5
|
作者
Zhang, Chihao [1 ,2 ,3 ]
Yang, Yang [4 ]
Zhou, Wei [4 ]
Zhang, Shihua [1 ,2 ,3 ]
机构
[1] Chinese Acad Sci, Acad Math & Syst Sci, RCSDS, NCMIS,CEMS, Beijing 100190, Peoples R China
[2] Univ Chinese Acad Sci, Sch Math Sci, Beijing 100049, Peoples R China
[3] Chinese Acad Sci, Ctr Excellence Anim Evolut & Genet, Kunming 650223, Yunnan, Peoples R China
[4] Yunnan Univ, Sch Software, Kunming 650504, Yunnan, Peoples R China
基金
中国国家自然科学基金;
关键词
Matrix decomposition; Bayes methods; Big Data; Principal component analysis; Distributed databases; Data mining; Clustering algorithms; Distributed algorithm; bayesian matrix decomposition; clustering; big data; data mining; FACTORIZATION; MODEL;
D O I
10.1109/TKDE.2020.3029582
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Matrix decomposition is one of the fundamental tools to discover knowledge from big data generated by modern applications. However, it is still inefficient or infeasible to process very big data using such a method in a single machine. Moreover, big data are often distributedly collected and stored on different machines. Thus, such data generally bear strong heterogeneous noise. It is essential and useful to develop distributed matrix decomposition for big data analytics. Such a method should scale up well, model the heterogeneous noise, and address the communication issue in a distributed system. To this end, we propose a distributed Bayesian matrix decomposition model (DBMD) for big data mining and clustering. Specifically, we adopt three strategies to implement the distributed computing including 1) the accelerated gradient descent, 2) the alternating direction method of multipliers (ADMM), and 3) the statistical inference. We investigate the theoretical convergence behaviors of these algorithms. To address the heterogeneity of the noise, we propose an optimal plug-in weighted average that reduces the variance of the estimation. Synthetic experiments validate our theoretical results, and real-world experiments show that our algorithms scale up well to big data and achieves superior or competing performance compared to two typical distributed methods including Scalable-NMF and scalable k-means++.
引用
收藏
页码:3701 / 3713
页数:13
相关论文
共 50 条
  • [41] An Improvement Approach for Reducing Dimensionality of Data with Matrix Decomposition in Data Mining
    Jamshidzadeh, Sasan
    Hosseinkhani, Javad
    [J]. INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND NETWORK SECURITY, 2016, 16 (12): : 11 - 14
  • [42] Distributed Stochastic Aware Random Forests - Efficient Data Mining for Big Data
    Assuncao, Joaquim
    Fernandes, Paulo
    Lopes, Lucelene
    Normey, Silvio
    [J]. 2013 IEEE INTERNATIONAL CONGRESS ON BIG DATA, 2013, : 425 - 426
  • [43] Green mining algorithm for big data based on random matrix
    [J]. Canwei, Wang (wangcanwei@sina.com), 1600, Science and Engineering Research Support Society (09):
  • [44] Parallel Clustering Optimization Algorithm Based on MapReduce in Big Data Mining
    Zhang, Huajie
    Song, Lei
    Zhang, Sen
    [J]. IAENG International Journal of Applied Mathematics, 2023, 53 (01):
  • [45] Image Information Mining: an Accelerated Bayesian Algorithm for Data Fusion of SAR Big Data
    Kevin, Alonso
    Mihai, Datcu
    [J]. 10TH EUROPEAN CONFERENCE ON SYNTHETIC APERTURE RADAR (EUSAR 2014), 2014,
  • [46] Bayesian Joint Matrix Decomposition for Data Integration with Heterogeneous Noise
    Zhang, Chihao
    Zhang, Shihua
    [J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2021, 43 (04) : 1184 - 1196
  • [47] User online behavior based on big data distributed clustering algorithm
    Wang, Yan
    [J]. INTERNATIONAL JOURNAL OF ADVANCED ROBOTIC SYSTEMS, 2020, 17 (02):
  • [48] Distributed evidential clustering toward time series with big data issue
    Gong, Chaoyu
    Su, Zhi-gang
    Wang, Pei-hong
    You, Yang
    [J]. EXPERT SYSTEMS WITH APPLICATIONS, 2022, 191
  • [49] A review on big data based parallel and distributed approaches of pattern mining
    Kumar, Sunil
    Mohbey, Krishna Kumar
    [J]. JOURNAL OF KING SAUD UNIVERSITY-COMPUTER AND INFORMATION SCIENCES, 2022, 34 (05) : 1639 - 1662
  • [50] Hadoop based Mining of Distributed Association Rules from Big Data
    Bouraoui, Marwa
    Bouzouita, Ines
    Touzi, Amel Grissa
    [J]. 2017 18TH INTERNATIONAL CONFERENCE ON SCIENCES AND TECHNIQUES OF AUTOMATIC CONTROL AND COMPUTER ENGINEERING (STA), 2017, : 185 - 190