1 paper · 1 filter
Julius Mayer, Mohamad Ballout, Serwan Jassim +2
Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal…