48 citations · 50 across the 5 of their papers we have counts for
1 paper · 1 filter
Jensen Gao, Bidipta Sarkar, Fei Xia +5
Recent advances in vision-language models (VLMs) have led to improved performance on tasks such as visual question answering and image captioning. Consequently, these models are no…