collaborators

7 papers

cs.CV2026

GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods

Sujay Belsare, Sudarshan Nikhil, Sushant Kumar +2

With the increasing development of Vision-Language Models, it becomes imperative that their predictions are readily explainable to relevant stakeholders. However, the field of expl…

cs.LG2026

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models

Ding Zhang, Runtao Zhou, Wenqing Zheng +3

Graph Language Models (GLMs) have become a promising direction for adapting Large Language Models (LLMs) to graph learning tasks. By transforming graph topology and node informatio…

cs.CV2026

Do Vision Language Models Need to Process Image Tokens?

Sambit Ghosh, R. Venkatesh Babu, Chirag Agarwal

Vision Language Models (VLMs) have achieved remarkable success by integrating visual encoders with large language models (LLMs). While VLMs process dense image tokens across deep t…

cs.CL2026

SynopticBench: Evaluating Vision-Language Models on Generating Weather Forecast Discussions of the Future

Timothy B. Higgins, Antonios Mamalakis, Chirag Agarwal

Recent advances in visual-language models (VLMs) have led to significant improvements in a plethora of complex multimodal tasks like image captioning, report generation, and visual…

cs.CL2025

Polarity-Aware Probing for Quantifying Latent Alignment in Language Models

Sabrina Sadiekh, Elena Ericheva, Chirag Agarwal

Advances in unsupervised probes such as Contrast-Consistent Search (CCS), which reveal latent beliefs without relying on token outputs, raise the question of whether these methods…

cs.AI2025

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar +6

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person…