3 papers
cs.CL2025
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
Julius Mayer, Mohamad Ballout, Serwan Jassim +2
Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal…
cs.CL2025
Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
Mohamad Ballout, Serwan Jassim, Elia Bruni
This paper presents a systematic evaluation of state-of-the-art multimodal large language models (MLLMs) on intuitive physics tasks using the GRASP and IntPhys 2 datasets. We asses…
cs.AI2024
Bidirectional Emergent Language in Situated Environments
Cornelius Wolff, Julius Mayer, Elia Bruni +1
Emergent language research has made significant progress in recent years, but still largely fails to explore how communication emerges in more complex and situated multi-agent syst…