A Differentially Private Kernel Two-Sample Test

被引：1

作者：

Raj, Anant ^{[1
]}

Law, Ho Chung Leon ^{[2
]}

Sejdinovic, Dino ^{[2
]}

Park, Mijung ^{[1
]}

机构：

[1] Max Planck Inst Intelligent Syst, Tubingen, Germany

[2] Univ Oxford, Dept Stat, Oxford, England

来源：

MACHINE LEARNING AND KNOWLEDGE DISCOVERY IN DATABASES, ECML PKDD 2019, PT I | 2020年 / 11906卷

关键词：

Differential privacy; Kernel two-sample test; NOISE;

D O I：

10.1007/978-3-030-46150-8_41

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Kernel two-sample testing is a useful statistical tool in determining whether data samples arise from different distributions without imposing any parametric assumptions on those distributions. However, raw data samples can expose sensitive information about individuals who participate in scientific studies, which makes the current tests vulnerable to privacy breaches. Hence, we design a new framework for kernel two-sample testing conforming to differential privacy constraints, in order to guarantee the privacy of subjects in the data. Unlike existing differentially private parametric tests that simply add noise to data, kernel-based testing imposes a challenge due to a complex dependence of test statistics on the raw data, as these statistics correspond to estimators of distances between representations of probability measures in Hilbert spaces. Our approach considers finite dimensional approximations to those representations. As a result, a simple chi-squared test is obtained, where a test statistic depends on a mean and covariance of empirical differences between the samples, which we perturb for a privacy guarantee. We investigate the utility of our framework in two realistic settings and conclude that our method requires only a relatively modest increase in sample size to achieve a similar level of power to the non-private tests in both settings.

引用

下载

页码：697 / 724

页数：28

共 50 条

[21] On a new multivariate two-sample test
Baringhaus, L
Franz, C
JOURNAL OF MULTIVARIATE ANALYSIS, 2004, 88 (01) : 190 - 206
[22] Revisiting the two-sample runs test
Baringhaus, Ludwig
Henze, Norbert
TEST, 2016, 25 (03) : 432 - 448
[23] The Bayesian two-sample t test
Gönen, M
Johnson, WO
Lu, YG
Westfall, PH
AMERICAN STATISTICIAN, 2005, 59 (03): : 252 - 257
[24] MMD Aggregated Two-Sample Test
Schrab, Antonin
Kim, Ilmun
Albert, Melisande
Laurent, Beatrice
Guedj, Benjamin
Gretton, Arthur
JOURNAL OF MACHINE LEARNING RESEARCH, 2023, 24
[25] Revisiting the two-sample runs test
Ludwig Baringhaus
Norbert Henze
TEST, 2016, 25 : 432 - 448
[26] Two-sample contamination model test
Milhaud, Xavier
Pommeret, Denys
Salhi, Yahia
Vandekerkhove, Pierre
BERNOULLI, 2024, 30 (01) : 170 - 197
[27] The score test for the two-sample occupancy model
Karavarsamis, N.
Guillera-Arroita, G.
Huggins, R. M.
Morgan, B. J. T.
AUSTRALIAN & NEW ZEALAND JOURNAL OF STATISTICS, 2020, 62 (01) : 94 - 115
[28] Two-Sample Hypothesis Test for Functional Data
Zhao, Jing
Feng, Sanying
Hu, Yuping
MATHEMATICS, 2022, 10 (21)
[29] A two-sample test when data are contaminated
Pommeret, Denys
STATISTICAL METHODS AND APPLICATIONS, 2013, 22 (04): : 501 - 516
[30] Two-sample test based on classification probability
Cai, Haiyan
Goggin, Bryan
Jiang, Qingtang
STATISTICAL ANALYSIS AND DATA MINING, 2020, 13 (01) : 5 - 13

← 1 2 3 4 5 →