Distributed statistical machine learning in adversarial settings: Byzantine gradient descent

被引：210

作者：

机构：

[1] Chen, Yudong

[2] Su, Lili

[3] Xu, Jiaming

来源：

| 1600年 / Association for Computing Machinery卷 / 01期

基金：

美国国家科学基金会;

关键词：

Learning systems - Parameter estimation - Machine learning;

D O I：

10.1145/3154503

中图分类号：

学科分类号：

摘要：

We consider the distributed statistical learning problem over decentralized systems that are prone to adversarial attacks. This setup arises in many practical applications, including Google's Federated Learning. Formally, we focus on a decentralized system that consists of a parameter server andm working machines; each working machine keeps N/m data samples, where N is the total number of samples. In each iteration, up to q of the m working machines suffer Byzantine faults-a faulty machine in the given iteration behaves arbitrarily badly against the system and has complete knowledge of the system. Additionally, the sets of faulty machines may be different across iterations. Our goal is to design robust algorithms such that the system can learn the underlying true parameter, which is of dimension d, despite the interruption of the Byzantine attacks. In this paper, based on the geometric median of means of the gradients, we propose a simple variant of the classical gradient descent method. We show that our method can tolerate q Byzantine failures up to 2(1 + ϵ )q ≤ m for an arbitrarily small but fixed constant ϵ > 0. The parameter estimate converges in O(log N) rounds with an estimation error on the order of max{ p dq/N, √ d/N}, which is larger than the minimax-optimal error rate √ d/N in the centralized and failure-free setting by at most a factor of √ q. The total computational complexity of our algorithm is of O((Nd/m) log N) at each working machine and O(md + kd log3 N) at the central server, and the total communication cost is of O(md log N). We further provide an application of our general results to the linear regression problem. A key challenge arises in the above problem is that Byzantine failures create arbitrary and unspecified dependency among the iterations and the aggregated gradients. To handle this issue in the analysis, we prove that the aggregated gradient, as a function of model parameter, converges uniformly to the true gradient function. © 2017 Association for Computing Machinery.

引用

共 50 条

[41] Leader Stochastic Gradient Descent for Distributed Training of Deep Learning Models
Teng, Yunfei
Gao, Wenbo
Chalus, Francois
Choromanska, Anna
Goldfarb, Donald
Weller, Adrian
ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 32 (NIPS 2019), 2019, 32
[42] Secure and Resilient Distributed Machine Learning Under Adversarial Environments
Zhang, Rui
Zhu, Quanyan
IEEE AEROSPACE AND ELECTRONIC SYSTEMS MAGAZINE, 2016, 31 (03) : 34 - 36
[43] Secure and Resilient Distributed Machine Learning Under Adversarial Environments
Zhang, Rui
Zhu, Quanyan
2015 18TH INTERNATIONAL CONFERENCE ON INFORMATION FUSION (FUSION), 2015, : 644 - 651
[44] Gradient descent machine learning regression for MHD flow: Metallurgy process
Priyadharshini, P.
Archana, M. Vanitha
Ahammad, N. Ameer
Raju, C. S. K.
Yook, Se-jin
Shah, Nehad Ali
INTERNATIONAL COMMUNICATIONS IN HEAT AND MASS TRANSFER, 2022, 138
[45] An Enhanced Optimization Scheme Based on Gradient Descent Methods for Machine Learning
Yi, Dokkyun
Ji, Sangmin
Bu, Sunyoung
SYMMETRY-BASEL, 2019, 11 (07):
[46] RECENT TRENDS IN STOCHASTIC GRADIENT DESCENT FOR MACHINE LEARNING AND BIG DATA
Newton, David
Pasupathy, Raghu
Yousefian, Farzad
2018 WINTER SIMULATION CONFERENCE (WSC), 2018, : 366 - 380
[47] Adaptive Distributed Learning with Byzantine Robustness : A Gradient-Projection-Based Method
Wang, Xinghan
Huang, Cheng
Ning, Jiahong
Yang, Tingting
Shen, Xuemin
IEEE CONFERENCE ON GLOBAL COMMUNICATIONS, GLOBECOM, 2023, : 7520 - 7525
[48] Accelerated Distributed Nesterov Gradient Descent
Qu, Guannan
Li, Na
IEEE TRANSACTIONS ON AUTOMATIC CONTROL, 2020, 65 (06) : 2566 - 2581
[49] Bayesian Distributed Stochastic Gradient Descent
Teng, Michael
Wood, Frank
ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 31 (NIPS 2018), 2018, 31
[50] Directed-Distributed Gradient Descent
Xi, Chenguang
Khan, Usman A.
2015 53RD ANNUAL ALLERTON CONFERENCE ON COMMUNICATION, CONTROL, AND COMPUTING (ALLERTON), 2015, : 1022 - 1026

← 1 2 3 4 5 →