C-sanitized: a privacy model for document redaction and sanitization
arXiv:1406.4285 · doi:10.1002/asi.23363
Abstract
Within the current context of Information Societies, large amounts of information are daily exchanged and/or released. The sensitive nature of much of this information causes a serious privacy threat when documents are uncontrollably made available to untrusted third parties. In such cases, appropriate data protection measures should be undertaken by the responsible organization, especially under the umbrella of current legislations on data privacy. To do so, human experts are usually requested to redact or sanitize document contents. To relieve this burdensome task, this paper presents a privacy model for document redaction/sanitization, which offers several advantages over other models available in the literature. Based on the well-established foundations of data semantics and the information theory, our model provides a framework to develop and implement automated and inherently semantic redaction/sanitization tools. Moreover, contrary to ad-hoc redaction methods, our proposal provides a priori privacy guarantees which can be intuitively defined according to current legislations on data privacy. Empirical tests performed within the context of several use cases illustrate the applicability of our model and its ability to mimic the reasoning of human sanitizers.
in Journal of the Association for Information Science and Technology, 2015
References in corpus (1)
Cited by in corpus (15)
- Toward sensitive document release with privacy guarantees
- Privacy-preserving data outsourcing in the cloud via semantic data splitting
- Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
- Is Your Model Sensitive? SPeDaC: A New Benchmark for Detecting and Classifying Sensitive Personal Data
- Enforcing transparent access to private content in social networks by means of automatic sanitization
- Deception for Cyber Defence: Challenges and Opportunities
- INRISCO: INcident monitoRing In Smart COmmunities
- On a Utilitarian Approach to Privacy Preserving Text Generation
- Privacy-Preserving Redaction of Diagnosis Data through Source Code Analysis
- Data Privacy in Multi-Cloud: An Enhanced Data Fragmentation Framework
- Silencing the Risk, Not the Whistle: A Semi-automated Text Sanitization Tool for Mitigating the Risk of Whistleblower Re-Identification
- A Differentially Private Text Perturbation Method Using a Regularized Mahalanobis Metric
- DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
- Sensitive Information Detection: Recursive Neural Networks for Encoding Context
- Ontology-based Access Control in Open Scenarios: Applications to Social Networks and the Cloud