Differential Privacy and Machine Learning: a Survey and Review
arXiv:1412.7584
Abstract
The objective of machine learning is to extract useful information from data, while privacy is preserved by concealing information. Thus it seems hard to reconcile these competing interests. However, they frequently must be balanced when mining sensitive data. For example, medical research represents an important application where it is necessary both to extract useful information and protect patient privacy. One way to resolve the conflict is to extract general characteristics of whole populations without disclosing the private information of individuals. In this paper, we consider differential privacy, one of the most popular and powerful definitions of privacy. We explore the interplay between machine learning and differential privacy, namely privacy-preserving machine learning algorithms and learning-based data release mechanisms. We also describe some theoretical results that address what can be learned differentially privately and upper bounds of loss functions for differentially private algorithms. Finally, we present some open questions, including how to incorporate public data, how to deal with missing data in private datasets, and whether, as the number of observed samples grows arbitrarily large, differentially private machine learning algorithms can be achieved at no cost to utility as compared to corresponding non-differentially private algorithms.
References in corpus (2)
Cited by in corpus (24)
- Differential Privacy-enabled Federated Learning for Sensitive Health Data
- When Machine Learning Meets Privacy: A Survey and Outlook
- Differential Privacy for Eye-Tracking Data
- Artificial Intelligence for Social Good: A Survey
- Differentially Private Synthetic Data: Applied Evaluations and Enhancements
- Differential Privacy in Natural Language Processing: The Story So Far
- DP-LSTM: Differential Privacy-inspired LSTM for Stock Prediction Using Financial News
- Privacy for All: Demystify Vulnerability Disparity of Differential Privacy against Membership Inference Attack
- Privacy-Preserving Blockchain Based Federated Learning with Differential Data Sharing
- Quantifying Membership Inference Vulnerability via Generalization Gap and Other Model Metrics
- AI and Ethics -- Operationalising Responsible AI
- Automated Anonymisation of Visual and Audio Data in Classroom Studies
- Privacy-preserving Active Learning on Sensitive Data for User Intent Classification
- Data collaboration analysis for distributed datasets
- Efficient hyperparameter optimization by way of PAC-Bayes bound minimization
- Know Your Model (KYM): Increasing Trust in AI and Machine Learning
- Differentially Private M-band Wavelet-Based Mechanisms in Machine Learning Environments
- Differential Privacy for Credit Risk Model
- Theoretical Model and Practical Considerations for Data Lineage Reconstruction
- Interpretable collaborative data analysis on distributed data
- Artificial Intelligence Narratives: An Objective Perspective on Current Developments
- Statistical Privacy Guarantees of Machine Learning Preprocessing Techniques
- A Framework for Behavior Privacy Preserving in Radio Frequency Signal
- Stability Enhanced Privacy and Applications in Private Stochastic Gradient Descent