1 paper
Jingyi Wang, Jianzhong Ju, Jian Luan +1
Recent advances in large vision-language models (VLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches…