Learning Differentially Private Recurrent Language Models
arXiv:1710.06963
Abstract
We demonstrate that it is possible to train large recurrent language models with user-level differential privacy guarantees with only a negligible cost in predictive accuracy. Our work builds on recent advances in the training of deep networks on user-partitioned data and privacy accounting for stochastic gradient descent. In particular, we add user-level privacy protection to the federated averaging algorithm, which makes "large step" updates from user-level data. Our work demonstrates that given a dataset with a sufficiently large number of users (a requirement easily met by even small internet-scale datasets), achieving differential privacy comes at the cost of increased computation, rather than in decreased utility as in most prior work. We find that our private LSTM language models are quantitatively and qualitatively similar to un-noised models when trained on a large dataset.
Camera-ready ICLR 2018 version, minor edits from previous
Cited by in corpus (198)
- On the Opportunities and Risks of Foundation Models
- A Survey on Federated Learning Systems: Vision, Hype and Reality for Data Privacy and Protection
- Towards Federated Learning at Scale: System Design
- How To Backdoor Federated Learning
- Can You Really Backdoor Federated Learning?
- Improving Federated Learning Personalization via Model Agnostic Meta Learning
- Personalized Federated Learning: A Meta-Learning Approach
- From Distributed Machine Learning to Federated Learning: A Survey
- FedPAQ: A Communication-Efficient Federated Learning Method with Periodic Averaging and Quantization
- Scalable Private Learning with PATE
- LEAF: A Benchmark for Federated Settings
- Extracting Training Data from Large Language Models
- Protection Against Reconstruction and Its Applications in Private Federated Learning
- Threats to Federated Learning: A Survey
- A Survey of Privacy Attacks in Machine Learning
- Personalized Federated Learning with Moreau Envelopes
- Inverting Gradients -- How easy is it to break privacy in federated learning?
- Federated Learning with Bayesian Differential Privacy
- DeepSight: Mitigating Backdoor Attacks in Federated Learning Through Deep Model Inspection
- Privacy Preserving Vertical Federated Learning for Tree-based Models
- A Field Guide to Federated Optimization
- Understanding Membership Inferences on Well-Generalized Learning Models
- How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
- Federated Learning and Differential Privacy: Software tools analysis, the Sherpa.ai FL framework and methodological guidelines for preserving data privacy
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Privacy-preserving Federated Learning for Residential Short Term Load Forecasting
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation
- Federated Forest
- Edge Intelligence: Architectures, Challenges, and Applications
- Attack of the Tails: Yes, You Really Can Backdoor Federated Learning
- Privacy in Deep Learning: A Survey
- Cronus: Robust and Heterogeneous Collaborative Learning with Black-Box Knowledge Transfer
- Analyzing Information Leakage of Updates to Natural Language Models
- No One Left Behind: Inclusive Federated Learning over Heterogeneous Devices
- Differentially Private Learning with Adaptive Clipping
- Systematic Evaluation of Privacy Risks of Machine Learning Models
- Differential Privacy Has Disparate Impact on Model Accuracy
- Large Language Models Can Be Strong Differentially Private Learners
- DQRE-SCnet: A novel hybrid approach for selecting users in Federated Learning with Deep-Q-Reinforcement Learning based on Spectral Clustering
- Real-World Image Datasets for Federated Learning
- Salvaging Federated Learning by Local Adaptation
- FedMix: Approximation of Mixup under Mean Augmented Federated Learning
- Federated Learning for Healthcare Informatics
- A Survey of Privacy Vulnerabilities of Mobile Device Sensors
- FedSAE: A Novel Self-Adaptive Federated Learning Framework in Heterogeneous Systems
- A Study of Face Obfuscation in ImageNet
- Differentially Private Learning Needs Better Features (or Much More Data)
- Label Leakage and Protection in Two-party Split Learning
- FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered Dropout
- Decentralized Deep Learning for Multi-Access Edge Computing: A Survey on Communication Efficiency and Trustworthiness
- Understanding Gradient Clipping in Private SGD: A Geometric Perspective
- Security and Privacy Issues in Deep Learning
- On Lightweight Privacy-Preserving Collaborative Learning for IoT Objects
- Boosting Privately: Privacy-Preserving Federated Extreme Boosting for Mobile Crowdsensing
- A Survey on Vulnerability of Federated Learning: A Learning Algorithm Perspective
- Enhanced Security and Privacy via Fragmented Federated Learning
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward Pass
- Differentially Private Fine-tuning of Language Models
- The OARF Benchmark Suite: Characterization and Implications for Federated Learning Systems
- Natural Language Understanding with Privacy-Preserving BERT
- Generative Models for Effective ML on Private, Decentralized Datasets
- Semi-Synchronous Federated Learning for Energy-Efficient Training and Accelerated Convergence in Cross-Silo Settings
- PPFL: Privacy-preserving Federated Learning with Trusted Execution Environments
- Privacy for Free: Communication-Efficient Learning with Differential Privacy Using Sketches
- Challenges of Privacy-Preserving Machine Learning in IoT
- Local Differential Privacy based Federated Learning for Internet of Things
- Privacy Accounting and Quality Control in the Sage Differentially Private ML Platform
- Federated Learning with Buffered Asynchronous Aggregation
- The Skellam Mechanism for Differentially Private Federated Learning
- Faster On-Device Training Using New Federated Momentum Algorithm
- LDP-FL: Practical Private Aggregation in Federated Learning with Local Differential Privacy
- Turbo-Aggregate: Breaking the Quadratic Aggregation Barrier in Secure Federated Learning
- GS-WGAN: A Gradient-Sanitized Approach for Learning Differentially Private Generators
- Federated Learning Meets Multi-objective Optimization
- Secure Federated Submodel Learning
- Asynchronous Federated Learning with Differential Privacy for Edge Intelligence
- Differentially-Private "Draw and Discard" Machine Learning
- Differentially Private Federated Learning on Heterogeneous Data
- A Scalable Approach for Privacy-Preserving Collaborative Machine Learning
- The Distributed Discrete Gaussian Mechanism for Federated Learning with Secure Aggregation
- Federated Reconstruction: Partially Local Federated Learning
- Enhancing the Privacy of Federated Learning with Sketching
- Federated Neural Architecture Search
- Training Data Leakage Analysis in Language Models
- Deep Learning with Gaussian Differential Privacy
- Hybrid Differentially Private Federated Learning on Vertically Partitioned Data
- Improving Deep Learning with Differential Privacy using Gradient Encoding and Denoising
- Practical and Private (Deep) Learning without Sampling or Shuffling
- Mitigating Membership Inference Attacks by Self-Distillation Through a Novel Ensemble Architecture
- Membership Inference Attack Susceptibility of Clinical Language Models
- D2P-Fed: Differentially Private Federated Learning With Efficient Communication
- Differentially Private Vertical Federated Clustering
- On the Outsized Importance of Learning Rates in Local Update Methods
- Federated Split Vision Transformer for COVID-19 CXR Diagnosis using Task-Agnostic Training
- Differentially Private Federated Learning for Resource-Constrained Internet of Things
- DPIS: An Enhanced Mechanism for Differentially Private SGD with Importance Sampling
- Federated Learning of N-gram Language Models
- Secure Weighted Aggregation for Federated Learning
- Learning discrete distributions: user vs item-level privacy
- Securing Federated Sensitive Topic Classification against Poisoning Attacks
- Shielding Collaborative Learning: Mitigating Poisoning Attacks through Client-Side Detection
- Fast-adapting and Privacy-preserving Federated Recommender System
- Federated Learning with Superquantile Aggregation for Heterogeneous Data
- Device Heterogeneity in Federated Learning: A Superquantile Approach
- Auditing Data Provenance in Text-Generation Models
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Federated Heavy Hitters Discovery with Differential Privacy
- Homogeneous Learning: Self-Attention Decentralized Deep Learning
- Optimal query complexity for private sequential learning against eavesdropping
- On the Impact of Device and Behavioral Heterogeneity in Federated Learning
- Group privacy for personalized federated learning
- Layer-wise Characterization of Latent Information Leakage in Federated Learning
- Information Leakage in Embedding Models
- Personalized Federated Learning with Gaussian Processes
- Efficient and Private Federated Learning with Partially Trainable Networks
- Network Shuffling: Privacy Amplification via Random Walks
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-Sharing
- FedSKETCH: Communication-Efficient and Private Federated Learning via Sketching
- Efficient Privacy-Preserving Stochastic Nonconvex Optimization
- Characterizing Impacts of Heterogeneity in Federated Learning upon Large-Scale Smartphone Data
- FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
- On the Practicality of Differential Privacy in Federated Learning by Tuning Iteration Times
- Instance-optimal Mean Estimation Under Differential Privacy
- Bayesian Differential Privacy for Machine Learning
- Element Level Differential Privacy: The Right Granularity of Privacy
- Weighted Distributed Differential Privacy ERM: Convex and Non-convex
- Differentially Private Federated Learning with Laplacian Smoothing
- Federated -Differential Privacy
- CaPC Learning: Confidential and Private Collaborative Learning
- Wide Network Learning with Differential Privacy
- On the Convergence and Calibration of Deep Learning with Differential Privacy
- PRECAD: Privacy-Preserving and Robust Federated Learning via Crypto-Aided Differential Privacy
- Auxo: Efficient Federated Learning via Scalable Client Clustering
- Fast Dimension Independent Private AdaGrad on Publicly Estimated Subspaces
- Evading Curse of Dimensionality in Unconstrained Private GLMs via Private Gradient Descent
- Concentrated Differentially Private and Utility Preserving Federated Learning
- Maximizing Uncertainty for Federated learning via Bayesian Optimisation-based Model Poisoning
- SoK: Training Machine Learning Models over Multiple Sources with Privacy Preservation
- Deep Learning Towards Mobile Applications
- Benchmarking Differential Privacy and Federated Learning for BERT Models
- Practical One-Shot Federated Learning for Cross-Silo Setting
- Voting-based Approaches For Differentially Private Federated Learning
- ESMFL: Efficient and Secure Models for Federated Learning
- Deep Learning with Label Differential Privacy
- Fairness-aware Differentially Private Collaborative Filtering
- Actor Critic with Differentially Private Critic
- Differentially Private Representation for NLP: Formal Guarantee and An Empirical Study on Privacy and Fairness
- Mitigating Sybil Attacks on Differential Privacy based Federated Learning
- TextHide: Tackling Data Privacy in Language Understanding Tasks
- Jointly Learning from Decentralized (Federated) and Centralized Data to Mitigate Distribution Shift
- Federated Learning in Adversarial Settings
- Differentially Private Deep Learning with Direct Feedback Alignment
- Tight and Robust Private Mean Estimation with Few Users
- Private Multi-Task Learning: Formulation and Applications to Federated Learning
- Synthetic data shuffling accelerates the convergence of federated learning under data heterogeneity
- Decentralized Differentially Private Segmentation with PATE
- Lossless Compression of Efficient Private Local Randomizers
- Source Inference Attacks in Federated Learning
- A Lightweight Privacy-Preserving Scheme Using Label-based Pixel Block Mixing for Image Classification in Deep Learning
- Policy-Based Federated Learning
- Personalised Federated Learning: A Combinational Approach
- Federated learning with differential privacy and an untrusted aggregator
- Private Language Model Adaptation for Speech Recognition
- Can Pretrained Language Models Derive Correct Semantics from Corrupt Subwords under Noise?
- DPlis: Boosting Utility of Differentially Private Deep Learning via Randomized Smoothing
- Citadel: Protecting Data Privacy and Model Confidentiality for Collaborative Learning with SGX
- P3SGD: Patient Privacy Preserving SGD for Regularizing Deep CNNs in Pathological Image Classification
- On the Privacy Risks of Deploying Recurrent Neural Networks in Machine Learning Models
- A Differentially Private Probabilistic Framework for Modeling the Variability Across Federated Datasets of Heterogeneous Multi-View Observations
- Make Text Unlearnable: Exploiting Effective Patterns to Protect Personal Data
- On Large-Cohort Training for Federated Learning
- Gradient Inversion with Generative Image Prior
- Towards Distributed Privacy-Preserving Prediction
- Federated Learning of User Verification Models Without Sharing Embeddings
- An Efficient DP-SGD Mechanism for Large Scale NLP Models
- Private Federated Learning Without a Trusted Server: Optimal Algorithms for Convex Losses
- Differential Privacy for Text Analytics via Natural Text Sanitization
- Fuzzi: A Three-Level Logic for Differential Privacy
- Revisiting Model-Agnostic Private Learning: Faster Rates and Active Learning
- Constrained Differentially Private Federated Learning for Low-bandwidth Devices
- PPT: A Privacy-Preserving Global Model Training Protocol for Federated Learning in P2P Networks
- DPack: Efficiency-Oriented Privacy Budget Scheduling
- Private Deep Learning with Teacher Ensembles
- Muffliato: Peer-to-Peer Privacy Amplification for Decentralized Optimization and Averaging
- From Noisy Fixed-Point Iterations to Private ADMM for Centralized and Federated Learning
- Multi-Frequency Federated Learning for Human Activity Recognition Using Head-Worn Sensors
- Universal Private Estimators
- Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language Models
- Federated Unbiased Learning to Rank
- Learning with User-Level Privacy
- A BIC-based Mixture Model Defense against Data Poisoning Attacks on Classifiers
- Gradient-Leakage Resilient Federated Learning
- Semi-Federated Learning
- Privacy enabled Financial Text Classification using Differential Privacy and Federated Learning
- Selective Differential Privacy for Language Modeling
- LINDT: Tackling Negative Federated Learning with Local Adaptation
- Multi-task Federated Edge Learning (MtFEEL) in Wireless Networks
- Federated Marginal Personalization for ASR Rescoring