3 citations · 4 across the 3 of their papers we have counts for
1 paper · 1 filter
Tharun Adithya Srikrishnan, Deval Shah, Timothy Hein +3
Large vision-language models (VLMs) enable joint processing of text and images. However, incorporating vision data significantly increases the prompt length, resulting in a longer…