3 papers
cs.CV2025
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
Monika Shah, Sudarshan Balaji, Somdeb Sarkhel +2
We can think of Visual Question Answering as a (multimodal) conversation between a human and an AI system. Here, we explore the sensitivity of Vision Language Models (VLMs) through…
cs.CV2025
On Explaining Visual Captioning with Hybrid Markov Logic Networks
Monika Shah, Somdeb Sarkhel, Deepak Venugopal
Deep Neural Networks (DNNs) have made tremendous progress in multimodal tasks such as image captioning. However, explaining/interpreting how these models integrate visual informati…
cs.CV2025
Disentangling Fine-Tuning from Pre-Training in Visual Captioning with Hybrid Markov Logic
Monika Shah, Somdeb Sarkhel, Deepak Venugopal
Multimodal systems have highly complex processing pipelines and are pretrained over large datasets before being fine-tuned for specific tasks such as visual captioning. However, it…