Generating synthetic personal health data using conditional generative adversarial networks combining with differential privacy

被引:9
|
作者
Sun, Chang [1 ,2 ]
van Soest, Johan [3 ,4 ]
Dumontier, Michel [1 ,2 ]
机构
[1] Maastricht Univ, Inst Data Sci, Fac Sci & Engn, Maastricht, Netherlands
[2] Maastricht Univ, Fac Sci & Engn, Dept Adv Comp Sci, Maastricht, Netherlands
[3] Maastricht Univ, Brightlands Inst Smart Soc, Fac Sci & Engn, Heerlen, Netherlands
[4] Maastricht Univ, GROW Sch Oncol & Reprod, Dept Radiat Oncol Maastro, Med Ctr, Maastricht, Netherlands
关键词
Synthetic data; Synthetic health data; Generative adversarial network; Data privacy; Health data sharing;
D O I
10.1016/j.jbi.2023.104404
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
A large amount of personal health data that is highly valuable to the scientific community is still not accessible or requires a lengthy request process due to privacy concerns and legal restrictions. As a solution, synthetic data has been studied and proposed to be a promising alternative to this issue. However, generating realistic and privacy-preserving synthetic personal health data retains challenges such as simulating the characteristics of the patients' data that are in the minority classes, capturing the relations among variables in imbalanced data and transferring them to the synthetic data, and preserving individual patients' privacy. In this paper, we propose a differentially private conditional Generative Adversarial Network model (DP-CGANS) consisting of data transformation, sampling, conditioning, and network training to generate realistic and privacy-preserving personal data. Our model distinguishes categorical and continuous variables and transforms them into latent space separately for better training performance. We tackle the unique challenges of generating synthetic patient data due to the special data characteristics of personal health data. For example, patients with a certain disease are typically the minority in the dataset and the relations among variables are crucial to be observed. Our model is structured with a conditional vector as an additional input to present the minority class in the imbalanced data and maximally capture the dependency between variables. Moreover, we inject statistical noise into the gradients in the networking training process of DP-CGANS to provide a differential privacy guarantee. We extensively evaluate our model with state-of-the-art generative models on personal socio-economic datasets and real-world personal health datasets in terms of statistical similarity, machine learning performance, and privacy measurement. We demonstrate that our model outperforms other comparable models, especially in capturing the dependence between variables. Finally, we present the balance between data utility and privacy in synthetic data generation considering the different data structures and characteristics of real-world personal health data such as imbalanced classes, abnormal distributions, and data sparsity.
引用
收藏
页数:14
相关论文
共 50 条
  • [1] Creation of Synthetic Data with Conditional Generative Adversarial Networks
    Vega-Marquez, Belen
    Rubio-Escudero, Cristina
    Riquelme, Jose C.
    Nepomuceno-Chamorro, Isabel
    [J]. 14TH INTERNATIONAL CONFERENCE ON SOFT COMPUTING MODELS IN INDUSTRIAL AND ENVIRONMENTAL APPLICATIONS (SOCO 2019), 2020, 950 : 231 - 240
  • [2] Generation of Synthetic Data with Conditional Generative Adversarial Networks
    Vega-Marquez, Belen
    Rubio-Escudero, Cristina
    Nepomuceno-Chamorro, Isabel
    [J]. LOGIC JOURNAL OF THE IGPL, 2022, 30 (02) : 252 - 262
  • [3] Generating Realistic Synthetic Traffic Data using Conditional Tabular Generative Adversarial Networks for Intelligent Transportation Systems
    Nigam, Archana
    Srivastava, Sanjay
    [J]. 2023 IEEE 26TH INTERNATIONAL CONFERENCE ON INTELLIGENT TRANSPORTATION SYSTEMS, ITSC, 2023, : 2881 - 2886
  • [4] Demand Side Data Generating Based on Conditional Generative Adversarial Networks
    Lan, Jian
    Guo, Qinglai
    Sun, Hongbin
    [J]. CLEANER ENERGY FOR CLEANER CITIES, 2018, 152 : 1188 - 1193
  • [5] On Generating Synthetic Histopathology Images Using Generative Adversarial Networks
    Carmody, Sean
    John, Deepu
    [J]. 2023 34TH IRISH SIGNALS AND SYSTEMS CONFERENCE, ISSC, 2023,
  • [6] Protecting Student Privacy with Synthetic Data from Generative Adversarial Networks
    Bautista, Peter
    Inventado, Paul Salvador
    [J]. ARTIFICIAL INTELLIGENCE IN EDUCATION (AIED 2021), PT II, 2021, 12749 : 66 - 70
  • [7] Generating Basic Unit Movements with Conditional Generative Adversarial Networks
    LUO Dingsheng
    NIE Mengxi
    WU Xihong
    [J]. Chinese Journal of Electronics, 2019, 28 (06) : 1099 - 1107
  • [8] Generating Basic Unit Movements with Conditional Generative Adversarial Networks
    Luo, Dingsheng
    Nie, Mengxi
    Wu, Xihong
    [J]. CHINESE JOURNAL OF ELECTRONICS, 2019, 28 (06) : 1099 - 1107
  • [9] Conditional Generative Adversarial Networks with Adversarial Attack and Defense for Generative Data Augmentation
    Baek, Francis
    Kim, Daeho
    Park, Somin
    Kim, Hyoungkwan
    Lee, SangHyun
    [J]. JOURNAL OF COMPUTING IN CIVIL ENGINEERING, 2022, 36 (03)
  • [10] Wasserstein Generative Adversarial Networks Based Differential Privacy Metaverse Data Sharing
    Liu H.
    Xu D.
    Tian Y.
    Peng C.
    Wu Z.
    Wang Z.
    [J]. IEEE Journal of Biomedical and Health Informatics, 2024, 28 (11) : 1 - 12