Scalable Privacy-Preserving Data Sharing Methodology for Genome-Wide Association Studies
arXiv:1401.5193 · doi:10.1016/j.jbi.2014.01.008
Abstract
The protection of privacy of individual-level information in genome-wide association study (GWAS) databases has been a major concern of researchers following the publication of "an attack" on GWAS data by Homer et al. (2008) Traditional statistical methods for confidentiality and privacy protection of statistical databases do not scale well to deal with GWAS data, especially in terms of guarantees regarding protection from linkage to external information. The more recent concept of differential privacy, introduced by the cryptographic community, is an approach that provides a rigorous definition of privacy with meaningful privacy guarantees in the presence of arbitrary external information, although the guarantees may come at a serious price in terms of data utility. Building on such notions, Uhler et al. (2013) proposed new methods to release aggregate GWAS data without compromising an individual's privacy. We extend the methods developed in Uhler et al. (2013) for releasing differentially-private -statistics by allowing for arbitrary number of cases and controls, and for releasing differentially-private allelic test statistics. We also provide a new interpretation by assuming the controls' data are known, which is a realistic assumption because some GWAS use publicly available data as controls. We assess the performance of the proposed methods through a risk-utility analysis on a real data set consisting of DNA samples collected by the Wellcome Trust Case Control Consortium and compare the methods with the differentially-private release mechanism proposed by Johnson and Shmatikov (2013).
28 pages, 2 figures, source code available upon request
References in corpus (1)
Cited by in corpus (25)
- Technical Privacy Metrics: a Systematic Survey
- An Economic Analysis of Privacy Protection and Statistical Accuracy as Social Choices
- Context-Aware Generative Adversarial Privacy
- Learning with Differential Privacy: Stability, Learnability and the Sufficiency and Necessity of ERM Principle
- Comparative Study of Differentially Private Data Synthesis Methods
- Artificial Intelligence for Social Good: A Survey
- Revisiting Differentially Private Hypothesis Tests for Categorical Data
- Generative Adversarial Privacy
- Differentially Private Chi-Squared Hypothesis Testing: Goodness of Fit and Independence Testing
- Confidentiality Protection in the 2020 US Census of Population and Housing
- Efficient differentially private learning improves drug sensitivity prediction
- SAFETY: Secure gwAs in Federated Environment Through a hYbrid solution with Intel SGX and Homomorphic Encryption
- A Primer on Private Statistics
- On-Average KL-Privacy and its equivalence to Generalization for Max-Entropy Mechanisms
- Priv'IT: Private and Sample Efficient Identity Testing
- Differentially Private False Discovery Rate Control
- Reproducibility and Transparency versus Privacy and Confidentiality: Reflections from a Data Editor
- A New Class of Private Chi-Square Tests
- Locally Private Hypothesis Testing
- Statistical Properties of Sanitized Results from Differentially Private Laplace Mechanism with Univariate Bounding Constraints
- INSPECTRE: Privately Estimating the Unseen
- Model-based Differentially Private Data Synthesis and Statistical Inference in Multiply Synthetic Differentially Private Data
- Privacy-preserving Stochastic Gradual Learning
- Privacy in the Genomic Era
- Near-Optimal Privacy-Utility Tradeoff in Genomic Studies Using Selective SNP Hiding