1 paper · 2 filters
Bhuvan Sachdeva, Karan Uppal, Abhinav Java +1
Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like depth estimation or object counting…