84 citations · 87 across the 2 of their papers we have counts for
1 paper · 1 filter
Boyuan Chen, Minghao Shao, Abdul Basit +2
As large language models (LLMs) grow more capable, they face growing vulnerability to sophisticated jailbreak attacks. While developers invest heavily in alignment finetuning and s…