Regularized Latent Semantic Indexing: A New Approach to Large-Scale Topic Modeling

被引：31

作者：

Wang, Quan ^{[1
]}

Xu, Jun ^{[2
]}

Li, Hang ^{[2
]}

Craswell, Nick ^{[3
]}

机构：

[1] Peking Univ, MOE Microsoft Key Lab Stat & Informat Technol, Beijing 100871, Peoples R China

[2] Microsoft Res Asia, Beijing, Peoples R China

[3] Microsoft Corp, Bellevue, WA 98004 USA

来源：

ACM TRANSACTIONS ON INFORMATION SYSTEMS | 2013年 / 31卷 / 01期

关键词：

Algorithms; Experimentation; Topic modeling; regularization; sparse methods; distributed learning; online learning;

D O I：

10.1145/2414782.2414787

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Topic modeling provides a powerful way to analyze the content of a collection of documents. It has become a popular tool in many research areas, such as text mining, information retrieval, natural language processing, and other related fields. In real-world applications, however, the usefulness of topic modeling is limited due to scalability issues. Scaling to larger document collections via parallelization is an active area of research, but most solutions require drastic steps, such as vastly reducing input vocabulary. In this article we introduce Regularized Latent Semantic Indexing (RLSI)-including a batch version and an online version, referred to as batch RLSI and online RLSI, respectively-to scale up topic modeling. Batch RLSI and online RLSI are as effective as existing topic modeling techniques and can scale to larger datasets without reducing input vocabulary. Moreover, online RLSI can be applied to stream data and can capture the dynamic evolution of topics. Both versions of RLSI formalize topic modeling as a problem of minimizing a quadratic loss function regularized by l(1) and/or l(2) norm. This formulation allows the learning process to be decomposed into multiple suboptimization problems which can be optimized in parallel, for example, via MapReduce. We particularly propose adopting l(1) norm on topics and l(2) norm on document representations to create a model with compact and readable topics and which is useful for retrieval. In learning, batch RLSI processes all the documents in the collection as a whole, while online RLSI processes the documents in the collection one by one. We also prove the convergence of the learning of online RLSI. Relevance ranking experiments on three TREC datasets show that batch RLSI and online RLSI perform better than LSI, PLSI, LDA, and NMF, and the improvements are sometimes statistically significant. Experiments on a Web dataset containing about 1.6 million documents and 7 million terms, demonstrate a similar boost in performance.

引用

页数：44

共 50 条

[1] Large-scale information retrieval with latent semantic indexing
Letsche, TA
Berry, MW
[J]. INFORMATION SCIENCES, 1997, 100 (1-4) : 105 - 137
[2] A Fast Approximate Algorithm for Large-Scale Latent Semantic Indexing
Zhang, Dell
Zhu, Zheng
[J]. 2008 THIRD INTERNATIONAL CONFERENCE ON DIGITAL INFORMATION MANAGEMENT, VOLS 1 AND 2, 2008, : 639 - 644
[3] Analyzing large-scale proteomics projects with latent semantic indexing
Klie, Sebastian
Martens, Lennart
Vizcaino, Juan Antonio
Cote, Richard
Jones, Phil
Apweiler, Rolf
Hinneburg, Alexander
Hermjakob, Henning
[J]. JOURNAL OF PROTEOME RESEARCH, 2008, 7 (01) : 182 - 191
[4] Regularized Latent Semantic Indexing
Wang, Quan
Xu, Jun
Li, Hang
Craswell, Nick
[J]. PROCEEDINGS OF THE 34TH INTERNATIONAL ACM SIGIR CONFERENCE ON RESEARCH AND DEVELOPMENT IN INFORMATION RETRIEVAL (SIGIR'11), 2011, : 685 - 694
[5] Hierarchy-regularized latent semantic indexing
Huang, Y
Yu, K
Schubert, M
Yu, SP
Tresp, V
Kriegel, HP
[J]. FIFTH IEEE INTERNATIONAL CONFERENCE ON DATA MINING, PROCEEDINGS, 2005, : 178 - 185
[6] Large-scale latent semantic analysis
Olney, Andrew McGregor
[J]. BEHAVIOR RESEARCH METHODS, 2011, 43 (02) : 414 - 423
[7] Large-scale latent semantic analysis
Andrew McGregor Olney
[J]. Behavior Research Methods, 2011, 43 : 414 - 423
[8] Automated Chinese Essay Scoring From Topic Perspective Using Regularized Latent Semantic Indexing
Hao, Shudong
Xu, Yanyan
Peng, Hengli
Su, Kaile
Ke, Dengfeng
[J]. 2014 22ND INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR), 2014, : 3092 - 3097
[9] Dynamic topic mapping using latent semantic indexing
Andres, F
Naito, M
[J]. Third International Conference on Information Technology and Applications, Vol 2, Proceedings, 2005, : 220 - 225
[10] Topic modeling for large-scale text data
Xi-ming Li
Ji-hong Ouyang
You Lu
[J]. Frontiers of Information Technology & Electronic Engineering, 2015, 16 : 457 - 465

← 1 2 3 4 5 →