1 paper · 1 filter
Roberta Rocca, Sami Boukortt, Geoff Keeling +1
Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to releva…