1 paper · 1 filter
Aishik Nagar, Shantanu Jaiswal, Cheston Tan
Vision-language models (VLMs) have shown impressive zero- and few-shot performance on real-world visual question answering (VQA) benchmarks, alluding to their capabilities as visua…