Privacy-Preserving Machine Learning: Methods, Challenges and Directions
arXiv:2108.04417
Abstract
Machine learning (ML) is increasingly being adopted in a wide variety of application domains. Usually, a well-performing ML model relies on a large volume of training data and high-powered computational resources. Such a need for and the use of huge volumes of data raise serious privacy concerns because of the potential risks of leakage of highly privacy-sensitive information; further, the evolving regulatory environments that increasingly restrict access to and use of privacy-sensitive data add significant challenges to fully benefiting from the power of ML for data-driven applications. A trained ML model may also be vulnerable to adversarial attacks such as membership, attribute, or property inference attacks and model inversion attacks. Hence, well-designed privacy-preserving ML (PPML) solutions are critically needed for many emerging applications. Increasingly, significant research efforts from both academia and industry can be seen in PPML areas that aim toward integrating privacy-preserving techniques into ML pipeline or specific algorithms, or designing various PPML architectures. In particular, existing PPML research cross-cut ML, systems and applications design, as well as security and privacy areas; hence, there is a critical need to understand state-of-the-art research, related challenges and a research roadmap for future research in PPML area. In this paper, we systematically review and summarize existing privacy-preserving approaches and propose a Phase, Guarantee, and Utility (PGU) triad based model to understand and guide the evaluation of various PPML solutions by decomposing their privacy-preserving functionalities. We discuss the unique characteristics and challenges of PPML and outline possible research directions that leverage as well as benefit multiple research communities such as ML, distributed systems, security and privacy.
References in corpus (20)
- Distilling the Knowledge in a Neural Network
- Language Models are Few-Shot Learners
- Private federated learning on vertically partitioned data via entity resolution and additively homomorphic encryption
- iDLG: Improved Deep Leakage from Gradients
- HybridAlpha: An Efficient Approach for Privacy-Preserving Federated Learning
- CryptoDL: Deep Neural Networks over Encrypted Data
- Threats to Federated Learning: A Survey
- FastSecAgg: Scalable Secure Aggregation for Privacy-Preserving Federated Learning
- Auditing Differentially Private Machine Learning: How Private is Private SGD?
- Privacy for Free: Communication-Efficient Learning with Differential Privacy Using Sketches
- Data Poisoning against Differentially-Private Learners: Attacks and Defenses
- Enhancing the Privacy of Federated Learning with Sketching
- CryptGPU: Fast Privacy-Preserving Machine Learning on the GPU
- Adaptive Histogram-Based Gradient Boosted Trees for Federated Learning
- FedSKETCH: Communication-Efficient and Private Federated Learning via Sketching
- Privacy-preserving Learning via Deep Net Pruning
- Mitigating Leakage in Federated Learning with Trusted Hardware
- PPCA: Privacy-preserving Principal Component Analysis Using Secure Multiparty Computation(MPC)
- Separation of Powers in Federated Learning
- Revisiting Secure Computation Using Functional Encryption: Opportunities and Research Directions