4 citations · 4 across the 1 of their papers we have counts for
3 papers
cs.LG2026
The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility
David Pape, Jonathan Evertz, Lea Schönherr
Progress in LLMs is increasingly measured through standardized benchmarks, where state-of-the-art improvements are often separated by fractions of a percentage point. At the same t…
cs.LG2026
No More, No Less: Task Alignment in Terminal Agents
Sina Mavali, David Pape, Jonathan Evertz +5
Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in…
cs.CR2025★ 4 cited
Chasing Shadows: Pitfalls in LLM Security Research
Jonathan Evertz, Niklas Risse, Nicolai Neuer +12
Large language models (LLMs) are increasingly prevalent in security research. Their unique characteristics, however, introduce challenges that undermine established paradigms of re…