4 papers
The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility
David Pape, Jonathan Evertz, Lea Schönherr
Progress in LLMs is increasingly measured through standardized benchmarks, where state-of-the-art improvements are often separated by fractions of a percentage point. At the same t…
No More, No Less: Task Alignment in Terminal Agents
Sina Mavali, David Pape, Jonathan Evertz +5
Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in…
Whispers in the Machine: Confidentiality in Agentic Systems
Jonathan Evertz, Merlin Chlosta, Lea Schönherr +1
Large language model (LLM)-based agents combine LLMs with external tools to automate tasks such as scheduling meetings, managing documents, or booking travel. While these integrati…
Chasing Shadows: Pitfalls in LLM Security Research
Jonathan Evertz, Niklas Risse, Nicolai Neuer +12
Large language models (LLMs) are increasingly prevalent in security research. Their unique characteristics, however, introduce challenges that undermine established paradigms of re…