1 paper · 1 filter
Siddharth Joshi, Besmira Nushi, Vidhisha Balachandran +4
Vision-language models (VLMs) are highly effective but often underperform on specialized tasks; for example, Llava-1.5 struggles with chart and diagram understanding due to scarce…