22 citations · 30 across the 8 of their papers we have counts for
1 paper · 2 filters
Karuna Bhaila, Aneesh Komanduri, Minh-Hao Van +1
Vision-Language Models (VLMs) have demonstrated immense capabilities in multi-modal understanding and inference tasks such as Visual Question Answering (VQA), which requires models…