1 citations · 1 across the 1 of their papers we have counts for
1 paper
Cuong Nhat Ha, Shima Asaadi, Sanjeev Kumar Karn +3
Vision-language models, while effective in general domains and showing strong performance in diverse multi-modal applications like visual question-answering (VQA), struggle to main…