3 papers
cs.AI2025
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
Jonathan Kutasov, Yuqi Sun, Paul Colognese +9
As Large Language Models (LLMs) are increasingly deployed as autonomous agents in complex and long horizon settings, it is critical to evaluate their ability to sabotage users by p…
cs.AI2025
The Elicitation Game: Evaluating Capability Elicitation Techniques
Felix Hofstätter, Teun van der Weij, Jayden Teoh +3
Capability evaluations are required to understand and regulate AI systems that may be deployed or further developed. Therefore, it is important that evaluations provide an accurate…
cs.LG2024
Extending Activation Steering to Broad Skills and Multiple Behaviours
Teun van der Weij, Massimo Poesio, Nandi Schoots
Current large language models have dangerous capabilities, which are likely to become more problematic in the future. Activation steering techniques can be used to reduce risks fro…