1 paper
Turhan Can Kargin, Wojciech Jasiński, Wojciech JasiÅski +5
Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicabili…