1 paper
Kaisheng Fan, Weizhe Zhang, Yishu Gao +2
Backdoored large language models (LLMs) exhibit attacker-specified behavior at inference time while retaining normal performance on benign inputs. Existing mitigations often requir…