5 papers
Can Vision Language Models Learn Intuitive Physics from Interaction?
Luca M. Schulze Buschoff, Konstantinos Voudouris, Can Demircan +1
Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can improve model performance on simple p…
Post-training makes large language models less human-like
Marcel Binz, Elif Akata, Abdullah Almaatouq +76
Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, w…
Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language Models
Luca M. Schulze Buschoff, Konstantinos Voudouris, Elif Akata +3
Pre-trained vision language models still fall short of human visual cognition. In an effort to improve visual cognition and align models with human behavior, we introduce visual st…
Centaur: a foundation model of human cognition
Marcel Binz, Elif Akata, Matthias Bethge +37
Establishing a unified theory of cognition has been a major goal of psychology. While there have been previous attempts to instantiate such theories by building computational model…
metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Alex Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff +1
Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the Open LLM Leaderboard aim to quantify these differences with several large benchmar…