4 papers
Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs
Kazuki Iwahana, Masaru Matsubayashi, Takuma Koyama +3
Backdoor attacks pose a serious threat to the safety and reliability of Large Language Models (LLMs), as they cause models to behave normally on clean inputs while producing attack…
Robust Backdoor Removal by Reconstructing Trigger-Activated Changes in Latent Representation
Kazuki Iwahana, Yusuke Yamasaki, Akira Ito +2
Backdoor attacks pose a critical threat to machine learning models, causing them to behave normally on clean data but misclassify poisoned data into a poisoned class. Existing defe…
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
Tomoya Yamashita, Akira Ito, Yuuki Yamanaka +3
As large language models (LLMs) are increasingly deployed across various applications, privacy and copyright concerns have heightened the need for more effective LLM unlearning tec…
Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
Tomoya Yamashita, Yuuki Yamanaka, Masanori Yamada +3
Machine Unlearning (MU) has recently attracted considerable attention as a solution to privacy and copyright issues in large language models (LLMs). Existing MU methods aim to remo…