10 citations · 21 across the 16 of their papers we have counts for
1 paper · 2 filters
Juntian Zhang, Chuanqi cheng, Yuhan Liu +3
Vision-language models (VLMs) achieve remarkable success in single-image tasks. However, real-world scenarios often involve intricate multi-image inputs, leading to a notable perfo…