2 papers
cs.AI2026
CPO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models
Swapnanil Mukherjee, Agyeya Negi, Tanuja Ganu +1
Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cro…
cs.CL2025
NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
Aritra Dutta, Swapnanil Mukherjee, Deepanway Ghosal +1
Commonsense visual-question answering often hinges on knowledge that is missing from the image or the question. Small vision-language models (sVLMs) such as ViLT, VisualBERT and FL…