A Latent Topic Model for Complete Entity Resolution

被引:0
|
作者
Shu, Liangcai [1 ]
Long, Bo [1 ]
Meng, Weiyi [1 ]
机构
[1] SUNY Binghamton, Dept Comp Sci, Binghamton, NY 13902 USA
关键词
DISTRIBUTIONS;
D O I
暂无
中图分类号
TP31 [计算机软件];
学科分类号
081202 ; 0835 ;
摘要
In bibliographies like DBLP and Citeseer, there are three kinds of entity-name problems that need to be solved. First, multiple entities share one name, which is called the name sharing problem. Second, one entity has different names, which is called the name variant problem. Third, multiple entities share multiple names, which is called the name mixing problem. We aim to solve these problems based on one model in this paper. We call this task complete entity resolution. Different from previous work, our work use global information based on data with two types of information, words and author names. We propose a generative latent topic model that involves both author names and words - the LDA-dual model, by extending the LDA (Latent Dirichlet Allocation) model. We also propose a method to obtain model parameters that is global information. Based on obtained model parameters, we propose two algorithms to solve the three problems mentioned above. Experimental results demonstrate the effectiveness and great potential of the proposed model and algorithms.
引用
收藏
页码:880 / 891
页数:12
相关论文
共 50 条
  • [1] A Latent Dirichlet Model for Unsupervised Entity Resolution
    Bhattacharya, Indrajit
    Getoor, Lise
    PROCEEDINGS OF THE SIXTH SIAM INTERNATIONAL CONFERENCE ON DATA MINING, 2006, : 47 - 58
  • [2] The Grouped Author-Topic Model for Unsupervised Entity Resolution
    Dai, Andrew M.
    Storkey, Amos J.
    ARTIFICIAL NEURAL NETWORKS AND MACHINE LEARNING - ICANN 2011, PT I, 2011, 6791 : 241 - 249
  • [3] Using Latent Topic Features for Named Entity Extraction in Search Queries
    Polifroni, Joe
    Mairesse, Francois
    12TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION 2011 (INTERSPEECH 2011), VOLS 1-5, 2011, : 2140 - 2143
  • [4] A Latent Topic Model for Linked Documents
    Guo, Zhen
    Zhu, Shenghuo
    Chi, Yun
    Zhang, Zhongfei
    Gong, Yihong
    PROCEEDINGS 32ND ANNUAL INTERNATIONAL ACM SIGIR CONFERENCE ON RESEARCH AND DEVELOPMENT IN INFORMATION RETRIEVAL, 2009, : 720 - 721
  • [5] Latent topic model for audio retrieval
    Hu, Pengfei
    Liu, Wenju
    Jiang, Wei
    Yang, Zhanlei
    PATTERN RECOGNITION, 2014, 47 (03) : 1138 - 1143
  • [6] Bayesian Latent Topic Clustering Model
    Wu, Meng-Sung
    Chien, Jen-Tzung
    INTERSPEECH 2008: 9TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION 2008, VOLS 1-5, 2008, : 2162 - 2165
  • [7] LATENT TOPIC MODEL FOR IMAGE ANNOTATION BY MODELING TOPIC CORRELATION
    Xu, Xing
    Shimada, Atsushi
    Taniguchi, Rin-ichiro
    2013 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO (ICME 2013), 2013,
  • [8] MixER: linear interpolation of latent space for entity resolution
    Wu, Huaiguang
    Li, Shuaichao
    COMPLEX & INTELLIGENT SYSTEMS, 2024, 10 (01) : 3 - 22
  • [9] MixER: linear interpolation of latent space for entity resolution
    Huaiguang Wu
    Shuaichao Li
    Complex & Intelligent Systems, 2024, 10 : 3 - 22
  • [10] Latent Topic Model for Indexing Arabic Documents
    Ayadi, Rami
    Maraoui, Mohsen
    Zrigui, Mounir
    INTERNATIONAL JOURNAL OF INFORMATION RETRIEVAL RESEARCH, 2014, 4 (02) : 57 - 72