From the 1 of 3 linked papers with an AI index.
1 paper · 1 filter
Euntae Kim, Soomin Han, Buru Chang
The paper introduces HarDBench, a benchmark that evaluates how vulnerable large language models are to jailbreak attacks when used as co-authors in draft-based writing, and propose…