1 paper · 1 filter
Ryle Goehausen, Marcus Sousa
Published evaluations of prompt-injection and jailbreak detectors for Large Language Models often suffer from two systematic weaknesses: per-dataset threshold tuning and undisclose…