2 papers
cs.MM2026
MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval
Aaryan Sharma, Vishak Prasad C, Virendra Singh +1
Vision-Language Models (VLMs) are highly effective in retrieving semantically relevant images. However, in practice, relevance alone is often insufficient. Systems must also achiev…
cs.CV2025
Enhancing Multi-Image Question Answering via Submodular Subset Selection
Aaryan Sharma, Shivansh Gupta, Samar Agarwal +2
Large multimodal models (LMMs) have achieved high performance in vision-language tasks involving single image but they struggle when presented with a collection of multiple images…