3 papers
cs.CR2026
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems
Leonhard Waibl, Felix Michalak, Hadrien Mariaccia
LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmar…
astro-ph.IM2026
Forecasting megaelectron-volt electron flux in the Earth's outer radiation belt using supervised machine learning algorithms and a timeseries foundation model
Rungployphan Kieokaew, Ryad Guezzi, François Ginisty +1
Accurate forecasting of megaelectron-volt (MeV) electrons in the outer Earth's radiation belt, which can pose significant risks to satellites, is essential for risk mitigation and…
cs.CR2025
The bitter lesson of misuse detection
Hadrien Mariaccia, Charbel-Raphaël Segerie, Diego Dorn
Prior work on jailbreak detection has established the importance of adversarial robustness for LLMs but has largely focused on the model ability to resist adversarial inputs and to…