2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Jaehoon Lee, Mingi Jung, Soohyuk Jang +3
Large Vision-Language Models (VLMs) achieve strong multimodal understanding capabilities by leveraging high-resolution visual inputs, but the resulting large number of visual token…