4 papers
Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +22
As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchma…
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
William Lugoloobi, Samuelle Marro, Jabez Magomere +2
As LLM-based agents increasingly browse the web on users' behalf, a natural question arises: can websites passively identify which underlying model powers an agent? Doing so would…
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
William Lugoloobi, Thomas Foster, William Bankes +1
Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remains challenging. We investigate whether the…
LLMs Encode How Difficult Problems Are
William Lugoloobi, Chris Russell
Large language models exhibit a puzzling inconsistency: they solve complex problems yet frequently fail on seemingly simpler ones. We investigate whether LLMs internally encode pro…