Arabic real time entity resolution using inverted indexing

被引:2
|
作者
Alian, Marwah [1 ,3 ]
Al-Naymat, Ghazi [2 ,3 ]
Ramadan, Banda [4 ]
机构
[1] Hashemite Univ, Zarqa, Jordan
[2] Ajman Univ, Ajman, U Arab Emirates
[3] Princess Sumaya Univ Technol, Amman, Jordan
[4] Prince Sultan Univ, Riyadh, Saudi Arabia
关键词
Arabic Entity Resolution; Similarity Aware Inverted Indexes; Similarity functions; Record pair comparison;
D O I
10.1007/s10579-020-09504-6
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Arabic datasets that have two or more records for the same world entity (i.e. person, object, etc.) make institutions suffer from low quality and degraded performance due to duplication in their Arabic datasets without having any mechanism for detecting these duplicates. The operation that distinguishes records for the same real-world entity is called Entity Resolution (ER). It is considered as a tool for linking records across databases as well as for matching query records with existing databases in real-time. Indexing is a major step in the ER process that aims at reducing the search space. Several indexing techniques are available for use with the ER process in general for English Databases. However, such techniques are not validated if they work well with other languages, such as Arabic. The Dynamic Similarity Aware Inverted Index (DySimII) is one of the indexing techniques that are utilized with dynamic databases to match query records in real time and is demonstrated to work well with English language. In this paper, we propose a framework-Arabic Real Time Entity Resolution (ARTER)-that uses DySimII with Arabic databases to perform real time ER. We also examine using different string similarity functions required for comparing records in the matching process for the aim of evaluating which similarity function is more suitable for comparing Arabic strings. A real-world Arabic database is used to conduct our experimental evaluation where two stemmers and three similarity functions are used to see the effect on DySimII with Arabic dataset. The results represent that matching accuracy is improved using Asem stemmer when the number of corrupted attributes is increased, also testing the three similarity functions show that using winkler similarity function provides better matching accuracy while N-gram provides better results when used with Asem stemmer.
引用
收藏
页码:921 / 941
页数:21
相关论文
共 50 条
  • [1] Arabic real time entity resolution using inverted indexing
    Marwah Alian
    Ghazi Al-Naymat
    Banda Ramadan
    Language Resources and Evaluation, 2020, 54 : 921 - 941
  • [2] Dynamic Sorted Neighborhood Indexing for Real-Time Entity Resolution
    Ramadan, Banda
    Christen, Peter
    Liang, Huizhi
    DATABASES THEORY AND APPLICATIONS, ADC 2014, 2014, 8506 : 1 - 12
  • [3] Dynamic Sorted Neighborhood Indexing for Real-Time Entity Resolution
    Ramadan, Banda
    Christen, Peter
    Liang, Huizhi
    Gayler, Ross W.
    ACM JOURNAL OF DATA AND INFORMATION QUALITY, 2015, 6 (04):
  • [4] Unsupervised learning blocking keys technique for indexing Arabic entity resolution
    Alian, Marwah
    Awajan, Arafat
    Ramadan, Bandan
    INTERNATIONAL JOURNAL OF SPEECH TECHNOLOGY, 2019, 22 (03) : 621 - 628
  • [5] Unsupervised learning blocking keys technique for indexing Arabic entity resolution
    Marwah Alian
    Arafat Awajan
    Bandan Ramadan
    International Journal of Speech Technology, 2019, 22 : 621 - 628
  • [6] Using Transliteration with Entity Resolution for Arabic Datasets
    Alian, Marwah
    Al-Naymat, Ghazi
    Ramadan, Banda
    2017 IEEE/ACS 14TH INTERNATIONAL CONFERENCE ON COMPUTER SYSTEMS AND APPLICATIONS (AICCSA), 2017, : 593 - 597
  • [7] A real time Named Entity Recognition system for Arabic text mining
    Al-Jumaily, Harith
    Martinez, Paloma
    Martinez-Fernandez, Jose L.
    Van der Goot, Erik
    LANGUAGE RESOURCES AND EVALUATION, 2012, 46 (04) : 543 - 563
  • [8] A real time Named Entity Recognition system for Arabic text mining
    Harith Al-Jumaily
    Paloma Martínez
    José L. Martínez-Fernández
    Erik Van der Goot
    Language Resources and Evaluation, 2012, 46 : 543 - 563
  • [9] Real-time Entity Resolution by Multiple Indices
    Zhu, Liang
    Cui, Rundong
    Ma, Qin
    Meng, Weiyi
    14TH INTERNATIONAL CONFERENCE ON COMPUTER SCIENCE AND EDUCATION (ICCSE 2019), 2019, : 1063 - 1068
  • [10] Unsupervised Blocking Key Selection for Real-Time Entity Resolution
    Ramadan, Banda
    Christen, Peter
    ADVANCES IN KNOWLEDGE DISCOVERY AND DATA MINING, PART II, 2015, 9078 : 574 - 585