1 paper · 1 filter
Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda +2
Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed t…