1 paper
Mikel Williams-Lekuona, Georgina Cosma
Vision transformers in vision-language models typically use the same amount of compute for every image, regardless of whether it is simple or complex. We propose ICAR (Image Comple…