Three-dimensional Entity Resolution with JedAI

被引:27
|
作者
Papadakis, George [1 ]
Mandilaras, George [1 ]
Gagliardelli, Luca [2 ]
Simonini, Giovanni [2 ]
Thanos, Emmanouil [3 ]
Giannakopoulos, George [4 ]
Bergamaschi, Sonia [2 ]
Palpanas, Themis [5 ,6 ]
Koubarakis, Manolis [1 ]
机构
[1] Natl & Kapodistrian Univ Athens, Athens, Greece
[2] Univ Modena & Reggio Emilia, Modena, Italy
[3] Katholieke Univ Leuven, Leuven, Belgium
[4] NCSR Demokritos, Paraskevi, Greece
[5] Univ Paris, Paris, France
[6] French Univ Inst IUF, Paris, France
基金
欧盟地平线“2020”;
关键词
Entity Resolution; Blocking; Matching; Clustering; Batch methods; Progressive methods; Massive parallelization; SIMILARITY JOINS; META-BLOCKING; LINKAGE;
D O I
10.1016/j.is.2020.101565
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Entity Resolution (ER) is the task of detecting different entity profiles that describe the same real-world objects. To facilitate its execution, we have developed JedAI, an open-source system that puts together a series of state-of-the-art ER techniques that have been proposed and examined independently, targeting parts of the ER end-to-end pipeline. This is a unique approach, as no other ER tool brings together so many established techniques. Instead, most ER tools merely convey a few techniques, those primarily developed by their creators. In addition to democratizing ER techniques, JedAI goes beyond the other ER tools by offering a series of unique characteristics: (i) It allows for building and benchmarking millions of ER pipelines. (ii) It is the only ER system that applies seamlessly to any combination of structured and/or semi-structured data. (iii) It constitutes the only ER system that runs seamlessly both on stand-alone computers and clusters of computers - through the parallel implementation of all algorithms in Apache Spark. (iv) It supports two different end-to-end workflows for carrying out batch ER (i.e., budget-agnostic), a schema-agnostic one based on blocks, and a schema-based one relying on similarity joins. (v) It adapts both end-to-end workflows to budget-aware (i.e., progressive) ER. We present in detail all features of JedAI, stressing the core characteristics that enhance its usability, and boost its versatility and effectiveness. We also compare it to the state-of-the-art in the field, qualitatively and quantitatively, demonstrating its state-of-the-art performance over a variety of large-scale datasets from different domains. The central repository of the JedAI's code base is here: https://github.com/scify/JedAIToolkit . A video demonstrating the JedAI's Web application is available here: https://www.youtube.com/watch?v=OJY1DUrUAe8. (C) 2020 Elsevier Ltd. All rights reserved.
引用
收藏
页数:17
相关论文
共 50 条
  • [31] High resolution three-dimensional numerical modelling of rockfalls
    Agliardi, F
    Crosta, GB
    INTERNATIONAL JOURNAL OF ROCK MECHANICS AND MINING SCIENCES, 2003, 40 (04) : 455 - 471
  • [32] Improvements in the mass resolution of the three-dimensional atom probe
    Sijbrandij, SJ
    Cerezo, A
    Godfrey, TJ
    Smith, GDW
    APPLIED SURFACE SCIENCE, 1996, 94-5 : 428 - 433
  • [33] Three-dimensional imaging of dislocations in a nanoparticle at atomic resolution
    Chen, Chien-Chun
    Zhu, Chun
    White, Edward R.
    Chiu, Chin-Yi
    Scott, M. C.
    Regan, B. C.
    Marks, Laurence D.
    Huang, Yu
    Miao, Jianwei
    NATURE, 2013, 496 (7443) : 74 - +
  • [34] Three-Dimensional Super Resolution Reconstruction by Integral Imaging
    Yang, Chen
    Wang, Jingang
    Stern, Adrian
    Gao, Shengkui
    Gurev, Viktor
    Javidi, Bahram
    JOURNAL OF DISPLAY TECHNOLOGY, 2015, 11 (11): : 947 - 952
  • [35] Three-dimensional Electrical Property Mapping with Nanometer Resolution
    Alekseev, Alexander
    Efimov, Anton
    Lu, Kangbo
    Loos, Joachim
    ADVANCED MATERIALS, 2009, 21 (48) : 4915 - +
  • [36] Three-dimensional resolution for circular synthetic aperture radar
    Moore, Linda J.
    Potter, Lee C.
    ALGORITHMS FOR SYNTHETIC APERTURE RADAR IMAGERY XIV, 2007, 6568
  • [37] High resolution three-dimensional prostate ultrasound imaging
    Li, Yinbo
    Patil, Abhay
    Hossack, John A.
    MEDICAL IMAGING 2006: ULTRASONIC IMAGING AND SIGNAL PROCESSING, 2006, 6147
  • [38] Measuring hydrophobic interactions with three-dimensional nanometer resolution
    Katan, Allard J.
    Oosterkamp, Tjerk H.
    JOURNAL OF PHYSICAL CHEMISTRY C, 2008, 112 (26): : 9769 - 9776
  • [39] Three-dimensional Optical-resolution Photoacoustic Microscopy
    Hu, Song
    Maslov, Konstantin
    Wang, Lihong V.
    JOVE-JOURNAL OF VISUALIZED EXPERIMENTS, 2011, (51):
  • [40] High resolution three-dimensional imaging with compress sensing
    Wang, Jingyi
    Ke, Jun
    OPTOELECTRONIC IMAGING AND MULTIMEDIA TECHNOLOGY IV, 2016, 10020