1 paper
Guangyu Shen, Siyuan Cheng, Xiangzhe Xu +4
Large Language Models (LLMs) can acquire deceptive behaviors through backdoor attacks, where the model executes prohibited actions whenever secret triggers appear in the input. Exi…