Differentially Private Learning with Adaptive Clipping
arXiv:1905.03871
Abstract
Existing approaches for training neural networks with user-level differential privacy (e.g., DP Federated Averaging) in federated learning (FL) settings involve bounding the contribution of each user's model update by clipping it to some constant value. However there is no good a priori setting of the clipping norm across tasks and learning settings: the update norm distribution depends on the model architecture and loss, the amount of data on each device, the client learning rate, and possibly various other parameters. We propose a method wherein instead of a fixed clipping norm, one clips to a value at a specified quantile of the update norm distribution, where the value at the quantile is itself estimated online, with differential privacy. The method tracks the quantile closely, uses a negligible amount of privacy budget, is compatible with other federated learning technologies such as compression and secure aggregation, and has a straightforward joint DP analysis with DP-FedAvg. Experiments demonstrate that adaptive clipping to the median update norm works well across a range of realistic federated learning tasks, sometimes outperforming even the best fixed clip chosen in hindsight, and without the need to tune any clipping hyperparameter.
Accepted to NeurIPS, 2021
References in corpus (5)
- Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification
- Practical Secure Aggregation for Federated Learning on User-Held Data
- Adaptive Federated Optimization
- A General Approach to Adding Differential Privacy to Iterative Training Procedures
- AdaCliP: Adaptive Clipping for Private SGD
Cited by in corpus (38)
- A Survey on Federated Learning Systems: Vision, Hype and Reality for Data Privacy and Protection
- How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
- Federated Learning for Healthcare Informatics
- Differentially Private Learning Needs Better Features (or Much More Data)
- Label Leakage and Protection in Two-party Split Learning
- Robust Federated Learning: The Case of Affine Distribution Shifts
- AdaCliP: Adaptive Clipping for Private SGD
- Understanding Gradient Clipping in Private SGD: A Geometric Perspective
- Attack-Resistant Federated Learning with Residual-based Reweighting
- DataLens: Scalable Privacy Preserving Training via Gradient Compression and Aggregation
- The Skellam Mechanism for Differentially Private Federated Learning
- Asynchronous Federated Learning with Differential Privacy for Edge Intelligence
- The Distributed Discrete Gaussian Mechanism for Federated Learning with Secure Aggregation
- Differentially Private Federated Learning on Heterogeneous Data
- Practical and Private (Deep) Learning without Sampling or Shuffling
- Understanding Unintended Memorization in Federated Learning
- Training Production Language Models without Memorizing User Data
- Fast-adapting and Privacy-preserving Federated Recommender System
- Instance-optimal Mean Estimation Under Differential Privacy
- Anonymizing Data for Privacy-Preserving Federated Learning
- Wide Network Learning with Differential Privacy
- Locally Differentially Private Analysis of Graph Statistics
- Privacy-preserving Non-negative Matrix Factorization with Outliers
- Removing Disparate Impact of Differentially Private Stochastic Gradient Descent on Model Accuracy
- Evading Curse of Dimensionality in Unconstrained Private GLMs via Private Gradient Descent
- Fast Dimension Independent Private AdaGrad on Publicly Estimated Subspaces
- Dynamic Differential-Privacy Preserving SGD
- ASFGNN: Automated Separated-Federated Graph Neural Network
- Differentially Private Deep Learning with Direct Feedback Alignment
- DP-REC: Private & Communication-Efficient Federated Learning
- On Large-Cohort Training for Federated Learning
- DPlis: Boosting Utility of Differentially Private Deep Learning via Randomized Smoothing
- Adaptive Differentially Private Empirical Risk Minimization
- An Efficient DP-SGD Mechanism for Large Scale NLP Models
- Universal Private Estimators
- Improving the Algorithm of Deep Learning with Differential Privacy
- Selective Differential Privacy for Language Modeling
- Efficient Hyperparameter Optimization for Differentially Private Deep Learning