High-Speed Rough Clustering for Very Large Document Collections

被引:8
|
作者
Kishida, Kazuaki [1 ]
机构
[1] Keio Univ, Sch Lib & Informat Sci, Minato Ku, Tokyo 1088345, Japan
关键词
CATEGORIZATION; SIMILARITY;
D O I
10.1002/asi.21311
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Document clustering is an important tool, but it is not yet widely used in practice probably because of its high computational complexity. This article explores techniques of high-speed rough clustering of documents, assuming that it is sometimes necessary to obtain a clustering result in a shorter time, although the result is just an approximate outline of document clusters. A promising approach for such clustering is to reduce the number of documents to be checked for generating cluster vectors in the leader follower clustering algorithm. Based on this idea, the present article proposes a modified Crouch algorithm and incomplete single-pass leader follower algorithm. Also, a two-stage grouping technique, in which the first stage attempts to decrease the number of documents to be processed in the second stage by applying a quick merging technique, is developed. An experiment using a part of the Reuters corpus RCV1 showed empirically that both the modified Crouch and the incomplete single-pass leader follower algorithms achieve clustering results more efficiently than the original methods, and also improved the effectiveness of clustering results. On the other hand, the two-stage grouping technique did not reduce the processing time in this experiment.
引用
收藏
页码:1092 / 1104
页数:13
相关论文
共 50 条
  • [1] Efficient clustering of very large document collections
    Dhillon, IS
    Fan, J
    Guan, YQ
    [J]. DATA MINING FOR SCIENTIFIC AND ENGINEERING APPLICATIONS, 2001, 2 : 357 - 381
  • [2] An efficient clustering approach for large document collections
    Han, B
    Kang, LS
    Song, HZ
    [J]. ADVANCED DATA MINING AND APPLICATIONS, PROCEEDINGS, 2005, 3584 : 240 - 247
  • [3] High-speed flames and DDT in very rough-walled channels
    Ciccarelli, Gaby
    Johansen, Craig
    Kellenberger, Mark
    [J]. COMBUSTION AND FLAME, 2013, 160 (01) : 204 - 211
  • [4] Trimaran configurations for high-speed very large ships
    Univ. degli Stud. Napoli Federico II, Italy
    [J]. Naval Architect, 2001, (AUGUST): : 28 - 30
  • [5] Trimaran configurations for high-speed very large ships
    Migali, A
    Miranda, S
    Pensa, C
    [J]. NAVAL ARCHITECT, 2001, : 28 - +
  • [6] Managing very large document collections using semantics
    GuoRen Wang
    HongJun Lu
    Ge Yu
    Bin YuBao
    [J]. Journal of Computer Science and Technology, 2003, 18 : 403 - 406
  • [7] Managing very large document collections using semantics
    Wang, GR
    Lu, HJ
    Yu, G
    Bao, YB
    [J]. JOURNAL OF COMPUTER SCIENCE AND TECHNOLOGY, 2003, 18 (03) : 403 - 406
  • [8] Multiplierless Design of High-Speed Very Large Constant Multiplications
    Aksoy, Levent
    Roy, Debapriya Basu
    Imran, Malik
    Pagliarini, Samuel
    [J]. 29TH ASIA AND SOUTH PACIFIC DESIGN AUTOMATION CONFERENCE, ASP-DAC 2024, 2024, : 957 - 962
  • [9] Topic modeling for mediated access to very large document collections
    Muresan, G
    Harper, DJ
    [J]. JOURNAL OF THE AMERICAN SOCIETY FOR INFORMATION SCIENCE AND TECHNOLOGY, 2004, 55 (10): : 892 - 910
  • [10] Possibilistic fuzzy co-clustering of large document collections
    Tjhi, William-Chandra
    Chen, Lihui
    [J]. PATTERN RECOGNITION, 2007, 40 (12) : 3452 - 3466