1 paper
Julius Mayer, Mohamad Ballout, Serwan Jassim +2
Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal…