55 citations · 87 across the 12 of their papers we have counts for
1 paper · 1 filter
Chenhui Zhang, Sherrie Wang
Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions. However, it remains unclear…