7 papers · 1 filter
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning
Puning Yang, Junchi Yu, Qizhou Wang +3
Mitigating sensitive and harmful outputs is fundamental to ensuring safe deployment of LLMs. Existing approaches typically follow two paradigms: Knowledge Deletion (KD), which eras…
Reflective Context Learning: Studying the Optimization Primitives of Context Space
Nikita Vassilyev, William Berrios, Ruowang Zhang +3
Generally capable agents must learn from experience in ways that generalize across tasks and environments. The fundamental problems of learning, including credit assignment, overfi…
LLM Unlearning with LLM Beliefs
Kemou Li, Qizhou Wang, Yue Wang +4
Large language models trained on vast corpora inherently risk memorizing sensitive or harmful content, which may later resurface in their outputs. Prevailing unlearning methods gen…
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
Fengpeng Li, Kemou Li, Qizhou Wang +2
Concept erasure helps stop diffusion models (DMs) from generating harmful content; but current methods face robustness retention trade off. Robustness means the model fine-tuned by…
GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
Yue Wang, Qizhou Wang, Feng Liu +4
Large language model (LLM) unlearning has demonstrated its essential role in removing privacy and copyright-related responses, crucial for their legal and safe applications. Howeve…
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
Puning Yang, Qizhou Wang, Zhuo Huang +3
Loss reweighting has shown significant benefits for machine unlearning with large language models (LLMs). However, their exact functionalities are left unclear and the optimal stra…