3 papers
cs.CL2026
Traces of Social Competence in Large Language Models
Tom Kouwenhoven, Michiel van der Meer, Max van Duijn
The False Belief Test (FBT) has been the main method for assessing Theory of Mind (ToM) and related socio-cognitive competencies. For Large Language Models (LLMs), the reliability…
cs.IR2026
Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?
Lennard C. Froma, Tom Kouwenhoven, Maaike H. T. de Boer +2
Much research on LLMs has focused on increasing benchmark performance. However, the evaluation of such models in real-world collaborative human-AI workflows has stayed behind. This…
cs.AI2025
Baba is LLM: Reasoning in a Game with Dynamic Rules
Fien van Wetten, Aske Plaat, Max van Duijn
Large language models (LLMs) are known to perform well on language tasks, but struggle with reasoning tasks. This paper explores the ability of LLMs to play the 2D puzzle game Baba…