4 citations · 4 across the 2 of their papers we have counts for
1 paper · 1 filter
William Hackett, Lewis Birch, Stefan Trawicki +2
Large Language Models (LLMs) guardrail systems are designed to protect against prompt injection and jailbreak attacks. However, they remain vulnerable to evasion techniques. We dem…