3 papers
cs.LG2026
No More, No Less: Task Alignment in Terminal Agents
Sina Mavali, David Pape, Jonathan Evertz +5
Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in…
cs.CL2026
Unknown Unknowns: Why Hidden Intentions in LLMs Evade Detection
Devansh Srivastav, David Pape, Lea Schönherr
LLMs are increasingly embedded in everyday decision-making, yet their outputs can encode subtle, unintended behaviours that shape user beliefs and actions. We refer to these covert…
cs.CR2025
Chasing Shadows: Pitfalls in LLM Security Research
Jonathan Evertz, Niklas Risse, Nicolai Neuer +12
Large language models (LLMs) are increasingly prevalent in security research. Their unique characteristics, however, introduce challenges that undermine established paradigms of re…