4 citations · 6 across the 6 of their papers we have counts for
1 paper · 1 filter
Yi Zeng, Weiyu Sun, Tran Ngoc Huynh +3
Safety backdoor attacks in large language models (LLMs) enable the stealthy triggering of unsafe behaviors while evading detection during normal interactions. The high dimensionali…