A Differentially Private Kernel Two-Sample Test

被引:1
|
作者
Raj, Anant [1 ]
Law, Ho Chung Leon [2 ]
Sejdinovic, Dino [2 ]
Park, Mijung [1 ]
机构
[1] Max Planck Inst Intelligent Syst, Tubingen, Germany
[2] Univ Oxford, Dept Stat, Oxford, England
关键词
Differential privacy; Kernel two-sample test; NOISE;
D O I
10.1007/978-3-030-46150-8_41
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Kernel two-sample testing is a useful statistical tool in determining whether data samples arise from different distributions without imposing any parametric assumptions on those distributions. However, raw data samples can expose sensitive information about individuals who participate in scientific studies, which makes the current tests vulnerable to privacy breaches. Hence, we design a new framework for kernel two-sample testing conforming to differential privacy constraints, in order to guarantee the privacy of subjects in the data. Unlike existing differentially private parametric tests that simply add noise to data, kernel-based testing imposes a challenge due to a complex dependence of test statistics on the raw data, as these statistics correspond to estimators of distances between representations of probability measures in Hilbert spaces. Our approach considers finite dimensional approximations to those representations. As a result, a simple chi-squared test is obtained, where a test statistic depends on a mean and covariance of empirical differences between the samples, which we perturb for a privacy guarantee. We investigate the utility of our framework in two realistic settings and conclude that our method requires only a relatively modest increase in sample size to achieve a similar level of power to the non-private tests in both settings.
引用
收藏
页码:697 / 724
页数:28
相关论文
共 50 条
  • [1] A Kernel Two-Sample Test
    Gretton, Arthur
    Borgwardt, Karsten M.
    Rasch, Malte J.
    Schoelkopf, Bernhard
    Smola, Alexander
    [J]. JOURNAL OF MACHINE LEARNING RESEARCH, 2012, 13 : 723 - 773
  • [2] A Kernel Two-Sample Test for Functional Data
    Wynne, George
    Duncan, Andrew B.
    [J]. JOURNAL OF MACHINE LEARNING RESEARCH, 2022, 23
  • [3] A Kernel Two-Sample Test for Functional Data
    Wynne, George
    Duncan, Andrew B.
    [J]. Journal of Machine Learning Research, 2022, 23 : 1 - 51
  • [4] Two-Sample Test with Kernel Projected Wasserstein Distance
    Wang, Jie
    Gao, Rui
    Xie, Yao
    [J]. INTERNATIONAL CONFERENCE ON ARTIFICIAL INTELLIGENCE AND STATISTICS, VOL 151, 2022, 151
  • [5] A permutation-free kernel two-sample test
    Shekhar, Shubhanshu
    Kim, Ilmun
    Ramdas, Aaditya
    [J]. ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 35 (NEURIPS 2022), 2022,
  • [6] The Kernel Two-Sample Test vs. Brain Decoding
    Olivetti, Emanuele
    Benozzo, Danilo
    Kia, Seyed Mostafa
    Ellero, Marta
    Hartmann, Thomas
    [J]. 2013 3RD INTERNATIONAL WORKSHOP ON PATTERN RECOGNITION IN NEUROIMAGING (PRNI 2013), 2013, : 128 - 131
  • [7] Sensor-level Maps with the Kernel Two-Sample Test
    Olivetti, Emanuele
    Kia, Seyed Mostafa
    Avesani, Paolo
    [J]. 2014 INTERNATIONAL WORKSHOP ON PATTERN RECOGNITION IN NEUROIMAGING, 2014,
  • [8] A TWO-SAMPLE TEST
    Moses, Lincoln E.
    [J]. PSYCHOMETRIKA, 1952, 17 (03) : 239 - 247
  • [9] A Bayesian Motivated Two-Sample Test Based on Kernel Density Estimates
    Merchant, Naveed
    Hart, Jeffrey D.
    [J]. ENTROPY, 2022, 24 (08)
  • [10] Network Traffic Fingerprinting Based on Approximated Kernel Two-Sample Test
    Kohout, Jan
    Pevny, Tomas
    [J]. IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, 2018, 13 (03) : 788 - 801