Arabic real time entity resolution using inverted indexing

被引:2
|
作者
Alian, Marwah [1 ,3 ]
Al-Naymat, Ghazi [2 ,3 ]
Ramadan, Banda [4 ]
机构
[1] Hashemite Univ, Zarqa, Jordan
[2] Ajman Univ, Ajman, U Arab Emirates
[3] Princess Sumaya Univ Technol, Amman, Jordan
[4] Prince Sultan Univ, Riyadh, Saudi Arabia
关键词
Arabic Entity Resolution; Similarity Aware Inverted Indexes; Similarity functions; Record pair comparison;
D O I
10.1007/s10579-020-09504-6
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Arabic datasets that have two or more records for the same world entity (i.e. person, object, etc.) make institutions suffer from low quality and degraded performance due to duplication in their Arabic datasets without having any mechanism for detecting these duplicates. The operation that distinguishes records for the same real-world entity is called Entity Resolution (ER). It is considered as a tool for linking records across databases as well as for matching query records with existing databases in real-time. Indexing is a major step in the ER process that aims at reducing the search space. Several indexing techniques are available for use with the ER process in general for English Databases. However, such techniques are not validated if they work well with other languages, such as Arabic. The Dynamic Similarity Aware Inverted Index (DySimII) is one of the indexing techniques that are utilized with dynamic databases to match query records in real time and is demonstrated to work well with English language. In this paper, we propose a framework-Arabic Real Time Entity Resolution (ARTER)-that uses DySimII with Arabic databases to perform real time ER. We also examine using different string similarity functions required for comparing records in the matching process for the aim of evaluating which similarity function is more suitable for comparing Arabic strings. A real-world Arabic database is used to conduct our experimental evaluation where two stemmers and three similarity functions are used to see the effect on DySimII with Arabic dataset. The results represent that matching accuracy is improved using Asem stemmer when the number of corrupted attributes is increased, also testing the three similarity functions show that using winkler similarity function provides better matching accuracy while N-gram provides better results when used with Asem stemmer.
引用
收藏
页码:921 / 941
页数:21
相关论文
共 50 条
  • [41] Improving Real Time Search Performance using Inverted Index Entries Invalidation Strategies
    Rissola, Esteban A.
    Tolosa, Gabriel H.
    JOURNAL OF COMPUTER SCIENCE & TECHNOLOGY, 2016, 16 (01): : 6 - 13
  • [42] Identification and real time control of an inverted pendulum using PI-PD controller
    Peker, Fuat
    Kaya, Ibrahim
    2017 21ST INTERNATIONAL CONFERENCE ON SYSTEM THEORY, CONTROL AND COMPUTING (ICSTCC), 2017, : 771 - 776
  • [43] Indexing Visual Features: Real-Time Loop Closure Detection Using a Tree Structure
    Liu, Yang
    Zhang, Hong
    2012 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION (ICRA), 2012, : 3613 - 3618
  • [44] Context-aware Approximate String Matching for Large-scale Real-time Entity Resolution
    Christen, Peter
    Gayler, Ross W.
    2015 IEEE INTERNATIONAL CONFERENCE ON DATA MINING WORKSHOP (ICDMW), 2015, : 211 - 217
  • [45] REAL-TIME SPOKEN ARABIC DIGIT RECOGNIZER
    ABDULLA, WH
    ABDULKARIM, MAH
    INTERNATIONAL JOURNAL OF ELECTRONICS, 1985, 59 (05) : 645 - 648
  • [46] Entity Resolution Using Inferred Relationships and Behavior
    Mugan, Jonathan
    Chari, Ranga
    Hitt, Laura
    McDermid, Eric
    Sowell, Marsha
    Qu, Yuan
    Coffman, Thayne
    2014 IEEE INTERNATIONAL CONFERENCE ON BIG DATA (BIG DATA), 2014, : 555 - 560
  • [47] Robust Entity Resolution using Random Graphs
    Galhotra, Sainyam
    Firmani, Donatella
    Saha, Barna
    Srivastava, Divesh
    SIGMOD'18: PROCEEDINGS OF THE 2018 INTERNATIONAL CONFERENCE ON MANAGEMENT OF DATA, 2018, : 3 - 18
  • [48] Entity Resolution Using Convolutional Neural Network
    Gottapu, Ram Deepak
    Dagli, Cihan
    Ali, Bharami
    COMPLEX ADAPTIVE SYSTEMS, 2016, 95 : 153 - 158
  • [49] Entity Resolution Acceleration using the Automata Processor
    Bo, Chunkun
    Wang, Ke
    Fox, Jeffrey J.
    Skadron, Kevin
    2016 IEEE INTERNATIONAL CONFERENCE ON BIG DATA (BIG DATA), 2016, : 311 - 318
  • [50] Real time adaptive fuzzy control of inverted pendulum
    Djurdjevic, PD
    42ND MIDWEST SYMPOSIUM ON CIRCUITS AND SYSTEMS, PROCEEDINGS, VOLS 1 AND 2, 1999, : 910 - 913