From the 1 of 4 linked papers with an AI index.
4 papers
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa +3
The paper introduces ParliamentBench, an open-source benchmark based on the Secret Hitler game, to evaluate large language model agents on deception, persuasion, and reasoning unde…
Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
Lars Benedikt Kaesberg, Tianyu Yang, Niklas Bauer +3
Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate models in a one-shot settin…
Evaluating Large Language Models in a Complex Hidden Role Game
Niklas Bauer
Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments. This work investigates the rea…
MALLM: Multi-Agent Large Language Models Framework
Jonas Becker, Lars Benedikt Kaesberg, Niklas Bauer +3
Multi-agent debate (MAD) has demonstrated the ability to augment collective intelligence by scaling test-time compute and leveraging expertise. Current frameworks for multi-agent d…