3 papers
cs.CL2026
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa +3
As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safe…
cs.AI2026
Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
Lars Benedikt Kaesberg, Tianyu Yang, Niklas Bauer +3
Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate models in a one-shot settin…
cs.MA2025
MALLM: Multi-Agent Large Language Models Framework
Jonas Becker, Lars Benedikt Kaesberg, Niklas Bauer +3
Multi-agent debate (MAD) has demonstrated the ability to augment collective intelligence by scaling test-time compute and leveraging expertise. Current frameworks for multi-agent d…