4 papers
Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce
Zeyuan Li, Lukas Petersson, Alessandro Acquisti +1
Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misalign…
Butter-Bench: Evaluating LLM Controlled Robots for Practical Intelligence
Callum Sharrock, Lukas Petersson, Hanna Petersson +4
We present Butter-Bench, a benchmark evaluating large language model (LLM) controlled robots for practical intelligence, defined as the ability to navigate the messiness of the phy…
Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models
Lukas Petersson, Axel Backlund, Axel Wennstöm +3
We introduce Blueprint-Bench, a benchmark designed to evaluate spatial reasoning capabilities in AI models through the task of converting apartment photographs into accurate 2D flo…
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Axel Backlund, Lukas Petersson
While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons. In…