1 paper · 1 filter
Tharun Adithya Srikrishnan, Deval Shah, Timothy Hein +3
Large vision-language models (VLMs) enable joint processing of text and images. However, incorporating vision data significantly increases the prompt length, resulting in a longer…