Three-dimensional Entity Resolution with JedAI

被引:27
|
作者
Papadakis, George [1 ]
Mandilaras, George [1 ]
Gagliardelli, Luca [2 ]
Simonini, Giovanni [2 ]
Thanos, Emmanouil [3 ]
Giannakopoulos, George [4 ]
Bergamaschi, Sonia [2 ]
Palpanas, Themis [5 ,6 ]
Koubarakis, Manolis [1 ]
机构
[1] Natl & Kapodistrian Univ Athens, Athens, Greece
[2] Univ Modena & Reggio Emilia, Modena, Italy
[3] Katholieke Univ Leuven, Leuven, Belgium
[4] NCSR Demokritos, Paraskevi, Greece
[5] Univ Paris, Paris, France
[6] French Univ Inst IUF, Paris, France
基金
欧盟地平线“2020”;
关键词
Entity Resolution; Blocking; Matching; Clustering; Batch methods; Progressive methods; Massive parallelization; SIMILARITY JOINS; META-BLOCKING; LINKAGE;
D O I
10.1016/j.is.2020.101565
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Entity Resolution (ER) is the task of detecting different entity profiles that describe the same real-world objects. To facilitate its execution, we have developed JedAI, an open-source system that puts together a series of state-of-the-art ER techniques that have been proposed and examined independently, targeting parts of the ER end-to-end pipeline. This is a unique approach, as no other ER tool brings together so many established techniques. Instead, most ER tools merely convey a few techniques, those primarily developed by their creators. In addition to democratizing ER techniques, JedAI goes beyond the other ER tools by offering a series of unique characteristics: (i) It allows for building and benchmarking millions of ER pipelines. (ii) It is the only ER system that applies seamlessly to any combination of structured and/or semi-structured data. (iii) It constitutes the only ER system that runs seamlessly both on stand-alone computers and clusters of computers - through the parallel implementation of all algorithms in Apache Spark. (iv) It supports two different end-to-end workflows for carrying out batch ER (i.e., budget-agnostic), a schema-agnostic one based on blocks, and a schema-based one relying on similarity joins. (v) It adapts both end-to-end workflows to budget-aware (i.e., progressive) ER. We present in detail all features of JedAI, stressing the core characteristics that enhance its usability, and boost its versatility and effectiveness. We also compare it to the state-of-the-art in the field, qualitatively and quantitatively, demonstrating its state-of-the-art performance over a variety of large-scale datasets from different domains. The central repository of the JedAI's code base is here: https://github.com/scify/JedAIToolkit . A video demonstrating the JedAI's Web application is available here: https://www.youtube.com/watch?v=OJY1DUrUAe8. (C) 2020 Elsevier Ltd. All rights reserved.
引用
收藏
页数:17
相关论文
共 50 条
  • [1] Reproducible experiments on Three-Dimensional Entity Resolution with JedAI
    Mandilaras, George
    Papadakis, George
    Gagliardelli, Luca
    Simonini, Giovanni
    Thanos, Emmanouil
    Giannakopoulos, George
    Bergamaschi, Sonia
    Palpanas, Themis
    Koubarakis, Manolis
    Lara-Clares, Alicia
    Farina, Antonio
    INFORMATION SYSTEMS, 2021, 102
  • [2] JedAI: The Force Behind Entity Resolution
    Papadakis, George
    Tsekouras, Leonidas
    Thanos, Emmanouil
    Giannakopoulos, George
    Palpanas, Themis
    Koubarakis, Manolis
    SEMANTIC WEB: ESWC 2017 SATELLITE EVENTS, 2017, 10577 : 161 - 166
  • [3] Three-dimensional Geospatial Interlinking with JedAI-spatial
    Papamichalopoulos, Marios
    Papadakis, George
    Mandilaras, George
    Siampou, Maria
    Mamoulis, Nikos
    Koubarakis, Manolis
    JOURNAL OF WEB SEMANTICS, 2024, 81
  • [4] Study on three-dimensional models of geological entity
    Shen, D.Y.
    Mao, S.J.
    Li, R.
    Ma, A.N.
    Jisuanji Gongcheng/Computer Engineering, 2001, 27 (03):
  • [5] Domain- and Structure-Agnostic End-to-End Entity Resolution with JedAI
    Papadakis, George
    Tsekouras, Leonidas
    Thanos, Emmanouil
    Giannakopoulos, George
    Palpanas, Themis
    Koubarakis, Manolis
    SIGMOD RECORD, 2019, 48 (04) : 30 - 36
  • [6] Three-dimensional subwavelength resolution with light
    Lewis, Aaron
    CURRENT OPINION IN STRUCTURAL BIOLOGY, 1991, 1 (06) : 1060 - 1064
  • [7] Depth resolution in three-dimensional images
    Son, Jung-Young
    Chernyshov, Oleksii
    Lee, Chun-Hae
    Park, Min-Chul
    Yano, Sumio
    JOURNAL OF THE OPTICAL SOCIETY OF AMERICA A-OPTICS IMAGE SCIENCE AND VISION, 2013, 30 (05) : 1030 - 1038
  • [8] A resolution measure for three-dimensional microscopy
    Chao, Jerry
    Ram, Sripad
    Abraham, Anish V.
    Ward, E. Sally
    Ober, Raimund J.
    OPTICS COMMUNICATIONS, 2009, 282 (09) : 1751 - 1761
  • [9] Three-dimensional imaging characteristics and depth resolution in digital holographic three-dimensional imaging spectrometry
    Obara, Masaki
    Yoshimori, Kyu
    INTERNATIONAL CONFERENCE ON PHOTONICS SOLUTIONS 2015, 2015, 9659
  • [10] Auto Dissection of Entity with Three-Dimensional Netwok based on FDTD
    Chen, L. L.
    Liao, C.
    Xia, X. Y.
    2010 ASIA-PACIFIC INTERNATIONAL SYMPOSIUM ON ELECTROMAGNETIC COMPATIBILITY & TECHNICAL EXHIBITION ON EMC RF/MICROWAVE MEASUREMENTS & INSTRUMENTATION, 2010, : 920 - 923