From the 1 of 4 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa +3
The paper introduces ParliamentBench, an open-source benchmark based on the Secret Hitler game, to evaluate large language model agents on deception, persuasion, and reasoning unde…
cs.CL2026
Evaluating Large Language Models in a Complex Hidden Role Game
Niklas Bauer
Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments. This work investigates the rea…