4 citations · 4 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility
David Pape, Jonathan Evertz, Lea Schönherr
Progress in LLMs is increasingly measured through standardized benchmarks, where state-of-the-art improvements are often separated by fractions of a percentage point. At the same t…
cs.LG2026
No More, No Less: Task Alignment in Terminal Agents
Sina Mavali, David Pape, Jonathan Evertz +5
Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in…