Enabling Synthetic Data adoption in regulated domains

被引:1
|
作者
Visani, Giorgio [1 ,2 ]
Graffi, Giacomo [2 ]
Alfero, Mattia [2 ]
Bagli, Enrico [2 ]
Chesani, Federico [1 ]
Capuzzo, Davide [2 ]
机构
[1] Univ Bologna, DISI Dept, Bologna, Italy
[2] CRIF SpA, R&D Dept, Bologna, Italy
关键词
synthetic data; benchmarks; goodness evaluation; data utility; privacy; finance; data sharing;
D O I
10.1109/DSAA54385.2022.10032356
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The switch from a Model-Centric to a Data-Centric mindset is putting emphasis on data and its quality rather than algorithms, bringing forward new challenges. In particular, the sensitive nature of the information in highly regulated scenarios needs to be accounted for. Specific approaches to address the privacy issue have been developed, as Privacy Enhancing Technologies. However, they frequently cause loss of information, putting forward a crucial trade-off among data quality and privacy. A clever way to bypass such a conundrum relies on Synthetic Data: data obtained from a generative process, learning the real data properties. Both Academia and Industry realized the importance of evaluating synthetic data quality: without all-round reliable metrics, the innovative data generation task has no proper objective function to maximize. Despite that, the topic remains under-explored. For this reason, we systematically catalog the important traits of synthetic data quality and privacy, and devise a specific methodology to test them. The result is DAISYnt (aDoption of Artificial Intelligence SYnthesis): a comprehensive suite of advanced tests, which sets a de facto standard for synthetic data evaluation. As a practical use-case, a variety of generative algorithms have been trained on real-world Credit Bureau Data. The best model has been assessed, using DAISYnt on the different synthetic replicas. Further potential uses, among others, entail auditing and fine-tuning of generative models or ensuring high quality of a given synthetic dataset. From a prescriptive viewpoint, eventually, DAISYnt may pave the way to synthetic data adoption in highly regulated domains, ranging from Finance to Healthcare, through Insurance and Education.
引用
收藏
页码:475 / 484
页数:10
相关论文
共 50 条
  • [1] Adoption of the Linked Data Best Practices in Different Topical Domains
    Schmachtenberg, Max
    Bizer, Christian
    Paulheim, Heiko
    [J]. SEMANTIC WEB - ISWC 2014, PT I, 2014, 8796 : 245 - 260
  • [2] Enabling Catalyst Adoption in SPARC
    Weirs, V. Gregory
    Raybourn, Elaine M.
    Milewicz, Reed
    Muollo, Killian
    Mauldin, Jeffrey A.
    Otahal, Thomas J.
    [J]. 2022 IEEE/ACM INTERNATIONAL WORKSHOP ON IN SITU INFRASTRUCTURES FOR ENABLING EXTREME-SCALE ANALYSIS AND VISUALIZATION (ISAV), 2022, : 20 - 25
  • [3] Enabling Near-Data Accelerators Adoption by Through Investigation of Datapath Solutions
    Santos, Paulo C.
    de Lima, Joao P. C.
    de Moura, Rafael F.
    Alves, Marco A. Z.
    Beck, Antonio C. S.
    Carro, Luigi
    [J]. INTERNATIONAL JOURNAL OF PARALLEL PROGRAMMING, 2021, 49 (02) : 237 - 252
  • [4] Enabling Near-Data Accelerators Adoption by Through Investigation of Datapath Solutions
    Paulo C. Santos
    João P. C. de Lima
    Rafael F. de Moura
    Marco A. Z. Alves
    Antonio C. S. Beck
    Luigi Carro
    [J]. International Journal of Parallel Programming, 2021, 49 : 237 - 252
  • [5] Synthetic data generation for digital twins: enabling production systems analysis in the absence of data
    Lopes, Paulo Victor
    Silveira, Leonardo
    Guimaraes Aquino, Roberto Douglas
    Ribeiro, Carlos Henrique
    Skoogh, Anders
    Verri, Filipe Alves Neto
    [J]. INTERNATIONAL JOURNAL OF COMPUTER INTEGRATED MANUFACTURING, 2024, 37 (10-11) : 1252 - 1269
  • [6] Enabling Wide Adoption of Hyperspectral Imaging
    Sharma, Neha
    [J]. MMSYS '21: PROCEEDINGS OF THE 2021 MULTIMEDIA SYSTEMS CONFERENCE, 2021, : 403 - 407
  • [7] Deep Aramaic: Towards a synthetic data paradigm enabling machine learning in epigraphy
    Aioanei, Andrei C.
    Hunziker-Rodewald, Regine R.
    Klein, Konstantin M.
    Michels, Dominik L.
    [J]. PLOS ONE, 2024, 19 (04):
  • [8] Enabling Robust Grammatical Error Correction in New Domains: Data Sets, Metrics, and Analyses
    Napoles C.
    Nădejde M.
    Tetreault J.
    [J]. Transactions of the Association for Computational Linguistics, 2019, 7 : 551 - 566
  • [9] Enabling Robust Grammatical Error Correction in New Domains: Data Sets, Metrics, and Analyses
    Napoles, Courtney
    Nadejde, Maria
    Tetreault, Joel
    [J]. TRANSACTIONS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, 2019, 7 : 551 - 566
  • [10] Securing Federated GANs: Enabling Synthetic Data Generation for Health Registry Consortiums
    Veeraragavan, Narasimha Raghavan
    Nygard, Jan F.
    [J]. 18TH INTERNATIONAL CONFERENCE ON AVAILABILITY, RELIABILITY & SECURITY, ARES 2023, 2023,