Linear Regression from Strategic Data Sources

被引:6
|
作者
Gast, Nicolas [1 ]
Ioannidis, Stratis [2 ]
Loiseau, Patrick [1 ,3 ]
Roussillon, Benjamin [1 ]
机构
[1] Univ Grenoble Alpes, LIG, Grenoble INP, Inria,CNRS, F-38000 Grenoble, France
[2] Northeastern Univ, Boston, MA 02120 USA
[3] MPI SWS, D-66123 Saarbrucken, Germany
关键词
Linear regression; Aitken theorem; Gauss-Markov theorem; strategic data sources; potential game; price of stability;
D O I
10.1145/3391436
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Linear regression is a fundamental building block of statistical data analysis. It amounts to estimating the parameters of a linear model that maps input features to corresponding outputs. In the classical setting where the precision of each data point is fixed, the famous Aitken/Gauss-Markov theorem in statistics states that generalized least squares (GLS) is a so-called "Best Linear Unbiased Estimator" (BLUE). In modern data science, however, one often faces strategic data sources; namely, individuals who incur a cost for providing high-precision data. For instance, this is the case for personal data, whose revelation may affect an individual's privacy-which can be modeled as a cost-or in applications such as recommender systems, where producing an accurate estimate entails effort. In this article, we study a setting in which features are public but individuals choose the precision of the outputs they reveal to an analyst. We assume that the analyst performs linear regression on this dataset, and individuals benefit from the outcome of this estimation. We model this scenario as a game where individuals minimize a cost composed of two components: (a) an (agent-specific) disclosure cost for providing high-precision data; and (b) a (global) estimation cost representing the inaccuracy in the linear model estimate. In this game, the linear model estimate is a public good that benefits all individuals. We establish that this game has a unique non-trivial Nash equilibrium. We study the efficiency of this equilibrium and we prove tight bounds on the price of stability for a large class of disclosure and estimation costs. Finally, we study the estimator accuracy achieved at equilibrium. We show that, in general, Aitken's theorem does not hold under strategic data sources, though it does hold if individuals have identical disclosure costs (up to a multiplicative factor). When individuals have non-identical costs, we derive a bound on the improvement of the equilibrium estimation cost that can be achieved by deviating from GLS, under mild assumptions on the disclosure cost functions.
引用
收藏
页数:24
相关论文
共 50 条
  • [41] Statistical Estimation with Strategic Data Sources in Competitive Settings
    Westenbroek, Tyler
    Dong, Roy
    Ratliff, Lillian J.
    Sastry, S. Shankar
    2017 IEEE 56TH ANNUAL CONFERENCE ON DECISION AND CONTROL (CDC), 2017,
  • [42] Local linear regression for generalized linear models with missing data
    Wang, CY
    Wang, SJ
    Gutierrez, RG
    Carroll, RJ
    ANNALS OF STATISTICS, 1998, 26 (03): : 1028 - 1050
  • [43] A linear programming approach to sparse linear regression with quantized data
    Cerone, V
    Fosson, S. M.
    Regruto, D.
    2019 AMERICAN CONTROL CONFERENCE (ACC), 2019, : 2990 - 2995
  • [44] Isolating and Examining Sources of Suppression and Multicollinearity in Multiple Linear Regression
    Beckstead, Jason W.
    MULTIVARIATE BEHAVIORAL RESEARCH, 2012, 47 (02) : 224 - 246
  • [45] Linear Regression from Uncertain Data and its Applications to Housing Price Prediction
    Gao, Shijian
    2020 3RD INTERNATIONAL CONFERENCE ON COMPUTER INFORMATION SCIENCE AND APPLICATION TECHNOLOGY (CISAT) 2020, 2020, 1634
  • [46] THE LINEAR CALIBRATION GRAPH AND ITS CONFIDENCE BANDS FROM REGRESSION ON TRANSFORMED DATA
    KURTZ, DA
    ROSENBERGER, JL
    TAMAYO, GJ
    ACS SYMPOSIUM SERIES, 1985, 284 : 133 - 165
  • [47] Remote Data Access and the Risk of Disclosure from Linear Regression: An Empirical Study
    Bleninger, Philipp
    Drechsler, Joerg
    Ronning, Gerd
    PRIVACY IN STATISTICAL DATABASES, 2010, 6344 : 220 - +
  • [48] Estimating monotonic rates from biological data using local linear regression
    Olito, Colin
    White, Craig R.
    Marshall, Dustin J.
    Barneche, Diego R.
    JOURNAL OF EXPERIMENTAL BIOLOGY, 2017, 220 (05): : 759 - 764
  • [49] Partial linear regression models for clustered data
    Chen, K
    Jin, ZZ
    JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2006, 101 (473) : 195 - 204
  • [50] Sequential linear regression with online standardized data
    Duarte, Kevin
    Monnez, Jean-Marie
    Albuisson, Eliane
    PLOS ONE, 2018, 13 (01):