Three-dimensional Entity Resolution with JedAI

被引:27
|
作者
Papadakis, George [1 ]
Mandilaras, George [1 ]
Gagliardelli, Luca [2 ]
Simonini, Giovanni [2 ]
Thanos, Emmanouil [3 ]
Giannakopoulos, George [4 ]
Bergamaschi, Sonia [2 ]
Palpanas, Themis [5 ,6 ]
Koubarakis, Manolis [1 ]
机构
[1] Natl & Kapodistrian Univ Athens, Athens, Greece
[2] Univ Modena & Reggio Emilia, Modena, Italy
[3] Katholieke Univ Leuven, Leuven, Belgium
[4] NCSR Demokritos, Paraskevi, Greece
[5] Univ Paris, Paris, France
[6] French Univ Inst IUF, Paris, France
基金
欧盟地平线“2020”;
关键词
Entity Resolution; Blocking; Matching; Clustering; Batch methods; Progressive methods; Massive parallelization; SIMILARITY JOINS; META-BLOCKING; LINKAGE;
D O I
10.1016/j.is.2020.101565
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Entity Resolution (ER) is the task of detecting different entity profiles that describe the same real-world objects. To facilitate its execution, we have developed JedAI, an open-source system that puts together a series of state-of-the-art ER techniques that have been proposed and examined independently, targeting parts of the ER end-to-end pipeline. This is a unique approach, as no other ER tool brings together so many established techniques. Instead, most ER tools merely convey a few techniques, those primarily developed by their creators. In addition to democratizing ER techniques, JedAI goes beyond the other ER tools by offering a series of unique characteristics: (i) It allows for building and benchmarking millions of ER pipelines. (ii) It is the only ER system that applies seamlessly to any combination of structured and/or semi-structured data. (iii) It constitutes the only ER system that runs seamlessly both on stand-alone computers and clusters of computers - through the parallel implementation of all algorithms in Apache Spark. (iv) It supports two different end-to-end workflows for carrying out batch ER (i.e., budget-agnostic), a schema-agnostic one based on blocks, and a schema-based one relying on similarity joins. (v) It adapts both end-to-end workflows to budget-aware (i.e., progressive) ER. We present in detail all features of JedAI, stressing the core characteristics that enhance its usability, and boost its versatility and effectiveness. We also compare it to the state-of-the-art in the field, qualitatively and quantitatively, demonstrating its state-of-the-art performance over a variety of large-scale datasets from different domains. The central repository of the JedAI's code base is here: https://github.com/scify/JedAIToolkit . A video demonstrating the JedAI's Web application is available here: https://www.youtube.com/watch?v=OJY1DUrUAe8. (C) 2020 Elsevier Ltd. All rights reserved.
引用
收藏
页数:17
相关论文
共 50 条
  • [21] Intramural coronary hematoma, a complex three-dimensional entity: Multimodality assessment
    Nunez-Gil, Ivan J.
    Echavarria-Pinto, Mauro
    Feltes, Gisela
    Fernandez-Ortiz, Antonio
    REVISTA PORTUGUESA DE CARDIOLOGIA, 2017, 36 (04)
  • [22] Left Atrium as a Dynamic Three-Dimensional Entity: Implications for Echocardiographic Assessment
    Badano, Luigi P.
    Nour, Angelica
    Muraru, Denisa
    REVISTA ESPANOLA DE CARDIOLOGIA, 2013, 66 (01): : 1 - 4
  • [23] A mechanistic macroscopic physical entity with a three-dimensional Hilbert space description
    Aerts, D
    Coecke, B
    DHooghe, B
    Valckenborgh, F
    HELVETICA PHYSICA ACTA, 1997, 70 (06): : 793 - 802
  • [24] Three-dimensional high resolution tomography for small objects
    Yamauchi, Yasushi
    Ikuta, Takashi
    Kishimoto, Naoki
    Nondestructive Testing and Evaluation, 1992, 7 (1 -4 pt 1) : 309 - 318
  • [25] A DETERMINATION OF THE THREE-DIMENSIONAL STRUCTURE OF TRICHOSANTHIN AT LOW RESOLUTION
    潘克桢
    张永茂
    林玉娟
    郑安
    陈元柱
    董贻诚
    陈世芝
    吴伸
    马星奇
    王耀萍
    伍伯牧
    窦士琦
    夏宗芗
    田庚元
    范肇昌
    倪朝周
    马益林
    孙孝先
    Science China Chemistry, 1982, (07) : 730 - 737
  • [26] High-resolution three-dimensional imaging of dislocations
    Barnard, J. S.
    Sharp, J.
    Tong, J. R.
    Midgley, P. A.
    SCIENCE, 2006, 313 (5785) : 319 - 319
  • [27] Three-dimensional, automated magnetic biomanipulation with subcellular resolution
    Schuerle, Simone
    Sakar, Mahmut Selman
    Meo, Alessandro
    Moeller, Jens
    Kratochvil, Bradley E.
    Chen, Christopher S.
    Vogel, Viola
    Nelson, Bradley J.
    2013 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION (ICRA), 2013, : 1452 - 1457
  • [28] Three-dimensional imaging of dislocations in a nanoparticle at atomic resolution
    Chien-Chun Chen
    Chun Zhu
    Edward R. White
    Chin-Yi Chiu
    M. C. Scott
    B. C. Regan
    Laurence D. Marks
    Yu Huang
    Jianwei Miao
    Nature, 2013, 496 : 74 - 77
  • [29] Three-Dimensional Structure of Brain Tissue at Submicrometer Resolution
    Saiga, Rino
    Mizutani, Ryuta
    Inomoto, Chie
    Takekoshi, Susumu
    Nakamura, Naoya
    Tsuboi, Akio
    Osawa, Motoki
    Arai, Makoto
    Oshima, Kenichi
    Itokawa, Masanari
    Uesugi, Kentaro
    Takeuchi, Akihisa
    Terada, Yasuko
    Suzuki, Yoshio
    XRM 2014: PROCEEDINGS OF THE 12TH INTERNATIONAL CONFERENCE ON X-RAY MICROSCOPY, 2016, 1696
  • [30] Recent advances in atomic resolution three-dimensional holography
    Daimon, Hiroshi
    Matsushita, Tomohiro
    Matsui, Fumihiko
    Hayashi, Koichi
    Wakabayashi, Yusuke
    ADVANCES IN PHYSICS-X, 2024, 9 (01):