1 paper · 1 filter
Aaryan Sharma, Shivansh Gupta, Samar Agarwal +2
Large multimodal models (LMMs) have achieved high performance in vision-language tasks involving single image but they struggle when presented with a collection of multiple images…