Incremental Factorization of Big Time Series Data with Blind Factor Approximation

被引：72

作者：

Chen, Dan ^{[1
]}

Tang, Yunbo ^{[1
]}

Zhang, Hao ^{[1
]}

Wang, Lizhe ^{[2
]}

Li, Xiaoli ^{[3
]}

机构：

[1] Wuhan Univ, Sch Comp Sci, Wuhan 430072, Peoples R China

[2] China Univ Geosci, Sch Comp Sci, Wuhan 430074, Peoples R China

[3] Beijing Normal Univ, Natl Key Lab Cognit Neurosci & Learning, Beijing 100875, Peoples R China

来源：

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING | 2021年 / 33卷 / 02期

基金：

中国国家自然科学基金;

关键词：

Big time series data; tensor factorization; blind factor approximation; parallel factor analysis; variational Bayesian inference; EEG; massively parallel computing; NONNEGATIVE MATRIX; MULTIWAY ANALYSIS; ALGORITHMS;

D O I：

10.1109/TKDE.2019.2931687

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Extracting the latent factors of big time series data is an important means to examine the dynamic complex systems under observation. These low-dimensional and "small" representations reveal the key insights to the overall mechanisms, which can otherwise be obscured by the notoriously high dimensionality and scale of big data as well as the enormously complicated interdependencies amongst data elements. However, grand challenges still remain: (1) to incrementally derive the multi-mode factors of the augmenting big data and (2) to achieve this goal under the circumstance of insufficient a priori knowledge. This study develops an incrementally parallel factorization solution (namely I-PARAFAC) for huge augmenting tensors (multi-way arrays) consisting of three phases over a cutting-edge GPU cluster: in the "giant-step" phase, a variational Bayesian inference (VBI) model estimates the distribution of the close neighborhood of each factor in a high confidence level without the need for a priori knowledge of the tensor or problem domain; in the "baby-step" phase, a massively parallel Fast-HALS algorithm (namely G-HALS) has been developed to derive the accurate subfactors of each subtensor on the basis of the initial factors; in the final fusion phase, I-PARAFAC fuses the known factors of the original tensor and those accurate subfactors of the "increment" to achieve the final full factors. Experimental results indicate that: (1) the VBI model enables a blind factor approximation, where the distribution of the close neighborhood of each final factor can be quickly derived (10 iterations for the test case). As a result, the model of a low time complexity significantly accelerates the derivation of the final accurate factors and lowers the risks of errors; (2) I-PARAFAC significantly outperforms even the latest high performance counterpart when handling augmenting tensors, e.g., the increased overhead is only proportional to the increment while the latter has to repeatedly factorize the whole tensor, and the overhead in fusing subfactors is always minimal; (3) I-PARAFAC can factorize a huge tensor (volume up to 500 TB over 50 nodes) as a whole with the capability several magnitudes higher than conventional methods, and the runtime is in the order of $\frac{1}{n}$1n to the number of compute nodes; (4) I-PARAFAC supports correct factorization-based analysis of a real 4-order EEG dataset captured from a variety of epilepsy patients. Overall, it should also be noted that counterpart methods have to derive the whole tensor from the scratch if the tensor is augmented in any dimension; as a contrast, the I-PARAFAC framework only needs to incrementally compute the full factors of the huge augmented tensor.

引用

页码：569 / 584

页数：16

共 50 条

[21] Incremental Learning in Time-series Data using Reinforcement Learning
Shuqair, Mustafa
Jimenez-Shahed, Joohi
Ghoraani, Behnaz
[J]. 2022 IEEE INTERNATIONAL CONFERENCE ON DATA MINING WORKSHOPS, ICDMW, 2022, : 868 - 875
[22] IncLSTM: Incremental Ensemble LSTM Model towards Time Series Data
Wang, Huiju
Li, Mengxuan
Yue, Xiao
[J]. COMPUTERS & ELECTRICAL ENGINEERING, 2021, 92
[23] Incremental Clustering for Time Series Data based on an Improved Leader Algorithm
Huynh Thi Thu Thuy
Duong Tuan Anh
Vo Thi Ngoc Chau
[J]. 2019 IEEE - RIVF INTERNATIONAL CONFERENCE ON COMPUTING AND COMMUNICATION TECHNOLOGIES (RIVF), 2019, : 13 - 18
[24] Assessment of Expert Interaction with Multivariate Time Series 'Big Data'
Adams, Susan Stevens
Haass, Michael J.
Matzen, Laura E.
King, Saskia
[J]. Foundations of Augmented Cognition: Neuroergonomics and Operational Neuroscience, Pt II, 2016, 9744 : 222 - 230
[25] TARDIS: Distributed Indexing Framework for Big Time Series Data
Zhang, Liang
Alghamdi, Noura
Eltabakh, Mohamed Y.
Rundensteiner, Elke A.
[J]. 2019 IEEE 35TH INTERNATIONAL CONFERENCE ON DATA ENGINEERING (ICDE 2019), 2019, : 1202 - 1213
[26] Smart Grid Time Series Big Data Processing System
Wang, Yuan
Yuan, Jun
Chen, Xiuming
Bao, Jianguo
[J]. 2015 IEEE ADVANCED INFORMATION TECHNOLOGY, ELECTRONIC AND AUTOMATION CONTROL CONFERENCE (IAEAC), 2015, : 393 - 400
[27] Big Data Analysis for Sensor Time-Series in Automation
Jirkovsky, Vaclav
Obitko, Marek
Novak, Petr
Kadera, Petr
[J]. 2014 IEEE EMERGING TECHNOLOGY AND FACTORY AUTOMATION (ETFA), 2014,
[28] Emulated order identification for models of big time series data
Wu, Brian
Drignei, Dorin
[J]. STATISTICAL ANALYSIS AND DATA MINING, 2021, 14 (02) : 201 - 212
[29] Compression Processing Estimation Method for Time Series Big Data
Miao Bei-bei
Jin Xue-bo
[J]. 2015 27TH CHINESE CONTROL AND DECISION CONFERENCE (CCDC), 2015, : 1807 - 1811
[30] THE POWER OF BIG DATA: HISTORICAL TIME SERIES ON GERMAN EDUCATION
Diebolt, Claude
Franzmann, Gabriele
Hippe, Ralph
Sensch, Jurgen
[J]. JOURNAL OF DEMOGRAPHIC ECONOMICS, 2017, 83 (03) : 329 - 376

← 1 2 3 4 5 →