1 citations · 1 across the 1 of their papers we have counts for
1 paper
Mingqian Feng, Yunlong Tang, Zeliang Zhang +1
Large Vision-Language Models (LVLMs) excel in integrating visual and linguistic contexts to produce detailed content, facilitating applications such as image captioning. However, u…