108 citations · 153 across the 9 of their papers we have counts for
1 paper · 2 filters
Zihan Weng, Lucas Gomez, Taylor Whittington Webb +1
Vision-Language Models (VLMs) have shown remarkable progress in visual understanding in recent years. Yet, they still lag behind human capabilities in specific visual tasks such as…