1 paper · 1 filter
Adithya TG, Adithya SK, Abhinav R Bharadwaj +2
Interacting and understanding with text heavy visual content with multiple images is a major challenge for traditional vision models. This paper is on enhancing vision models' capa…