1 paper · 1 filter
Zhiyuan Fan, Yumeng Wang, Sandeep Polisetty +1
Large Vision Language Models (LVLMs) excel in various vision-language tasks. Yet, their robustness to visual variations in position, scale, orientation, and context that objects in…