Practical and Private (Deep) Learning without Sampling or Shuffling
arXiv:2103.00039
Abstract
We consider training models with differential privacy (DP) using mini-batch gradients. The existing state-of-the-art, Differentially Private Stochastic Gradient Descent (DP-SGD), requires privacy amplification by sampling or shuffling to obtain the best privacy/accuracy/computation trade-offs. Unfortunately, the precise requirements on exact sampling and shuffling can be hard to obtain in important practical scenarios, particularly federated learning (FL). We design and analyze a DP variant of Follow-The-Regularized-Leader (DP-FTRL) that compares favorably (both theoretically and empirically) to amplified DP-SGD, while allowing for much more flexible data access patterns. DP-FTRL does not use any form of privacy amplification. The code is available at https://github.com/google-research/federated/tree/master/dp_ftrl and https://github.com/google-research/DP-FTRL .
References in corpus (8)
- Towards Federated Learning at Scale: System Design
- Adaptive Bound Optimization for Online Convex Optimization
- Auditing Differentially Private Machine Learning: How Private is Private SGD?
- Encode, Shuffle, Analyze Privacy Revisited: Formalizations and Empirical Evaluation
- Tempered Sigmoid Activations for Deep Learning with Differential Privacy
- Understanding Unintended Memorization in Federated Learning
- Stability of Stochastic Gradient Descent on Nonsmooth Convex Losses
- Training Production Language Models without Memorizing User Data
Cited by in corpus (10)
- How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
- FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
- Federated Learning with Buffered Asynchronous Aggregation
- The Skellam Mechanism for Differentially Private Federated Learning
- The Distributed Discrete Gaussian Mechanism for Federated Learning with Secure Aggregation
- SoK: Machine Learning Governance
- The More, the Better? A Study on Collaborative Machine Learning for DGA Detection
- On the Convergence and Calibration of Deep Learning with Differential Privacy
- DP-SGD vs PATE: Which Has Less Disparate Impact on Model Accuracy?
- DPack: Efficiency-Oriented Privacy Budget Scheduling