4 citations · 5 across the 2 of their papers we have counts for
1 paper · 1 filter
David Wan, Jaemin Cho, Elias Stengel-Eskin +1
Highlighting particularly relevant regions of an image can improve the performance of vision-language models (VLMs) on various vision-language (VL) tasks by guiding the model to at…