ForgetMe: Evaluating Selective Forgetting in Generative Models
arXiv:2504.12574 · doi:10.1016/j.engappai.2025.112087
Abstract
The widespread adoption of diffusion models in image generation has increased the demand for privacy-compliant unlearning. However, due to the high-dimensional nature and complex feature representations of diffusion models, achieving selective unlearning remains challenging, as existing methods struggle to remove sensitive information while preserving the consistency of non-sensitive regions. To address this, we propose an Automatic Dataset Creation Framework based on prompt-based layered editing and training-free local feature removal, constructing the ForgetMe dataset and introducing the Entangled evaluation metric. The Entangled metric quantifies unlearning effectiveness by assessing the similarity and consistency between the target and background regions and supports both paired (Entangled-D) and unpaired (Entangled-S) image data, enabling unsupervised evaluation. The ForgetMe dataset encompasses a diverse set of real and synthetic scenarios, including CUB-200-2011 (Birds), Stanford-Dogs, ImageNet, and a synthetic cat dataset. We apply LoRA fine-tuning on Stable Diffusion to achieve selective unlearning on this dataset and validate the effectiveness of both the ForgetMe dataset and the Entangled metric, establishing them as benchmarks for selective unlearning. Our work provides a scalable and adaptable solution for advancing privacy-preserving generative AI.
References in corpus (18)
- LoRA: Low-Rank Adaptation of Large Language Models
- On the Opportunities and Risks of Foundation Models
- Machine Unlearning: Solutions and Challenges
- Selective Amnesia: A Continual Learning Approach to Forgetting in Deep Generative Models
- Inst-Inpaint: Instructing to Remove Objects with Diffusion Models
- Large Language Model Unlearning
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
- Fast Model Debias with Machine Unlearning
- Yuan: Yielding Unblemished Aesthetics Through A Unified Network for Visual Imperfections Removal in Generated Images
- Ferrari: Federated Feature Unlearning via Optimizing Feature Sensitivity
- Score Forgetting Distillation: A Swift, Data-Free Method for Machine Unlearning in Diffusion Models
- Machine Unlearning in Large Language Models
- CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
- On Large Language Model Continual Unlearning
- Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models
- Evaluating Deep Unlearning in Large Language Models
- Offset Unlearning for Large Language Models
- Machine Unlearning for Traditional Models and Large Language Models: A Short Survey