1 paper · 1 filter
Becky Mashaido, Tapadhir Das
Prompt injection attacks pose significant risks to language model safety, yet existing defenses are typically evaluated using classification performance. We show that high detectio…