Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
Ben Slater, Matteo G. Mecattaf, Lucy G. Cheke +2
Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autono…
cs.CL2026
Visuospatial Perspective Taking in Multimodal Language Models
Jonathan Prunty, Seraphina Zhang, Patrick Quinn +3
As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks l…
cs.CL2024
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
Lorenzo Pacchiardi, Marko Tesic, Lucy G. Cheke +1
The integrity of AI benchmarks is fundamental to accurately assess the capabilities of AI systems. The internal validity of these benchmarks - i.e., making sure they are free from…