4 citations · 4 across the 13 of their papers we have counts for
1 paper · 1 filter
Hanjiang Hu, Alexander Robey, Changliu Liu
Large language models (LLMs) are shown to be vulnerable to jailbreaking attacks where adversarial prompts are designed to elicit harmful responses. While existing defenses effectiv…