2 papers
cs.CR2025
The bitter lesson of misuse detection
Hadrien Mariaccia, Charbel-Raphaël Segerie, Diego Dorn
Prior work on jailbreak detection has established the importance of adversarial robustness for LLMs but has largely focused on the model ability to resist adversarial inputs and to…
cs.CR2024
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
Diego Dorn, Alexandre Variengien, Charbel-Raphaël Segerie +1
Input-output safeguards are used to detect anomalies in the traces produced by Large Language Models (LLMs) systems. These detectors are at the core of diverse safety-critical appl…