1 paper · 1 filter
Zihan Weng, Lucas Gomez, Taylor Whittington Webb +1
Vision-Language Models (VLMs) have shown remarkable progress in visual understanding in recent years. Yet, they still lag behind human capabilities in specific visual tasks such as…