1 paper · 1 filter
Seunghee Koh, Sunghyun Baek, Youngdong Kim +1
Unlearning in large language models (LLMs) has emerged as a promising safeguard against adversarial behaviors. When the forgetting loss is applied uniformly without considering tok…