4 papers · 1 filter
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
Julius Mayer, Mohamad Ballout, Serwan Jassim +2
Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal…
Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
Mohamad Ballout, Serwan Jassim, Elia Bruni
This paper presents a systematic evaluation of state-of-the-art multimodal large language models (MLLMs) on intuitive physics tasks using the GRASP and IntPhys 2 datasets. We asses…
Interpretability of Language Models via Task Spaces
Lucas Weber, Jaap Jumelet, Elia Bruni +1
The usual way to interpret language models (LMs) is to test their performance on different benchmarks and subsequently infer their internal processes. In this paper, we present an…
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
Serwan Jassim, Mario Holubar, Annika Richter +3
This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This…