2 papers
cs.CV2026
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
Amartya Bhattacharya
Vision-language models (VLMs) excel at image-text retrieval yet persistently fail at compositional reasoning, distinguishing captions that share the same words but differ in relati…
cs.LG2025
Can Out-of-Domain data help to Learn Domain-Specific Prompts for Multimodal Misinformation Detection?
Amartya Bhattacharya, Debarshi Brahma, Suraj Nagaje Mahadev +3
Spread of fake news using out-of-context images and captions has become widespread in this era of information overload. Since fake news can belong to different domains like politic…