1 paper · 1 filter
Matthew Stephenson, Matthew Sidji, Benoît Ronval
In this paper, we propose the use of the popular word-based board game Codenames as a suitable benchmark for evaluating the reasoning capabilities of Large Language Models (LLMs).…