1 paper · 1 filter
Yifan Luo, Zhennan Zhou, Meitan Wang +1
In this paper, we investigate the safety mechanisms of instruction fine-tuned large language models (LLMs). We discover that re-weighting MLP neurons can significantly compromise a…