2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Jiabao Ji, Bairu Hou, Alexander Robey +5
Aligned large language models (LLMs) are vulnerable to jailbreaking attacks, which bypass the safeguards of targeted LLMs and fool them into generating objectionable content. While…