Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods
Sujay Belsare, Sudarshan Nikhil, Sushant Kumar +2
With the increasing development of Vision-Language Models, it becomes imperative that their predictions are readily explainable to relevant stakeholders. However, the field of expl…
cs.CV2026
Do Vision Language Models Need to Process Image Tokens?
Sambit Ghosh, R. Venkatesh Babu, Chirag Agarwal
Vision Language Models (VLMs) have achieved remarkable success by integrating visual encoders with large language models (LLMs). While VLMs process dense image tokens across deep t…
cs.CV2025
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
Ashish Seth, Dinesh Manocha, Chirag Agarwal
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance in complex multimodal tasks. However, these models still suffer from hallucinations, particularly when…