Scalable Discovery and Continuous Inventory of Personal Data at Rest in Cloud Native Systems
arXiv:2209.10412 · doi:10.1007/978-3-031-20984-0_36
Abstract
Cloud native systems are processing large amounts of personal data through numerous and possibly multi-paradigmatic data stores (e.g., relational and non-relational databases). From a privacy engineering perspective, a core challenge is to keep track of all exact locations, where personal data is being stored, as required by regulatory frameworks such as the European General Data Protection Regulation. In this paper, we present Teiresias, comprising i) a workflow pattern for scalable discovery of personal data at rest, and ii) a cloud native system architecture and open source prototype implementation of said workflow pattern. To this end, we enable a continuous inventory of personal data featuring transparency and accountability following DevOps/DevPrivOps practices. In particular, we scope version-controlled Infrastructure as Code definitions, cloud-based storages, and how to integrate the process into CI/CD pipelines. Thereafter, we provide iii) a comparative performance evaluation demonstrating both appropriate execution times for real-world settings, and a promising personal data detection accuracy outperforming existing proprietary tools in public clouds.
Preprint of 2022-09-09 before final copy-editing of an accepted peer-reviewed paper to appear in the Proceedings of the 20th International Conference on Service-Oriented Computing ICSOC 2022
References in corpus (5)
- Continuous Integration, Delivery and Deployment: A Systematic Review on Approaches, Tools, Challenges and Practices
- TILT: A GDPR-Aligned Transparency Information Language and Toolkit for Practical Privacy Engineering
- The GDPR Enforcement Fines at Glance
- Breyer case of the Court of Justice of the European Union: IP addresses and the personal data definition
- Cloud Native Privacy Engineering through DevPrivOps