Document Classification Using Nonnegative Matrix Factorization and Underapproximation

被引:13
|
作者
Berry, Michael W. [1 ]
Gillis, Nicolas [1 ]
Glineur, Francois [1 ]
机构
[1] Univ Tennessee, Dept Elect Engn & Comp Sci, Knoxville, TN 37996 USA
关键词
ALGORITHMS;
D O I
10.1109/ISCAS.2009.5118379
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
In this study, we use nonnegative matrix factorization (NMF) and nonnegative matrix underapproximation (NMU) approaches to generate feature vectors that can be used to cluster Aviation Safety Reporting System (ASRS) documents obtained from the Distributed National ASAP Archive (DNAA). By preserving nonnegativity, both the NMF and NMU facilitate a sum-of-parts representation of the underlying term usage patterns in the ASRS document collection. Both the training and test sets of ASRS documents are parsed and then factored by both algorithms to produce a reduced-rank representations of the entire document space. The resulting feature and coefficient matrix factors are used to cluster ASRS documents so that the (known) associated anomalies of training documents are directly mapped to the feature vectors. Dominant features of test documents are then used to generate anomaly relevance scores for those documents. We demonstrate that the approximate solution obtained by NMU using Lagrangrian duality can lead to a better sum-of-parts representation and document classification accuracy.
引用
收藏
页码:2782 / 2785
页数:4
相关论文
共 50 条
  • [21] Gene selection and cancer classification using Monte Carlo and nonnegative matrix factorization
    Chen, Jing
    Ma, Qin
    Hu, Xiaoyan
    Zhang, Miao
    Qin, Dongdong
    Lu, Xiaoquan
    [J]. RSC ADVANCES, 2016, 6 (46) : 39652 - 39656
  • [22] Nonnegative Matrix Factorization
    不详
    [J]. IEEE CONTROL SYSTEMS MAGAZINE, 2021, 41 (03): : 102 - 102
  • [23] Distributional Clustering Using Nonnegative Matrix Factorization
    Zhu, Zhenfeng
    Ye, Yangdong
    [J]. PROCEEDINGS OF THE 10TH WORLD CONGRESS ON INTELLIGENT CONTROL AND AUTOMATION (WCICA 2012), 2012, : 4705 - 4711
  • [24] Nonnegative Matrix Factorization
    SAIBABA, A. R. V. I. N. D. K.
    [J]. SIAM REVIEW, 2022, 64 (02) : 510 - 511
  • [25] Automated Graph Regularized Projective Nonnegative Matrix Factorization for Document Clustering
    Pei, Xiaobing
    Wu, Tao
    Chen, Chuanbo
    [J]. IEEE TRANSACTIONS ON CYBERNETICS, 2014, 44 (10) : 1821 - 1831
  • [26] Anomaly detection using nonnegative matrix factorization
    Allan, Edward G.
    Horvath, Michael R.
    Kopek, Christopher V.
    Lamb, Brian T.
    Whaples, Thomas S.
    Berry, Michael W.
    [J]. SURVEY OF TEXT MINING II: CLUSTERING, CLASSIFICATION, AND RETRIEVAL, 2008, : 203 - +
  • [27] Community discovery using nonnegative matrix factorization
    Wang, Fei
    Li, Tao
    Wang, Xin
    Zhu, Shenghuo
    Ding, Chris
    [J]. DATA MINING AND KNOWLEDGE DISCOVERY, 2011, 22 (03) : 493 - 521
  • [28] Using underapproximations for sparse nonnegative matrix factorization
    Gillis, Nicolas
    Glineur, Francois
    [J]. PATTERN RECOGNITION, 2010, 43 (04) : 1676 - 1687
  • [29] Community discovery using nonnegative matrix factorization
    Fei Wang
    Tao Li
    Xin Wang
    Shenghuo Zhu
    Chris Ding
    [J]. Data Mining and Knowledge Discovery, 2011, 22 : 493 - 521
  • [30] Structure constrained nonnegative matrix factorization for pattern clustering and classification
    Lu, Na
    Miao, Hongyu
    [J]. NEUROCOMPUTING, 2016, 171 : 400 - 411