1 paper
Ruoxiang Huang, Xindian Ma, Rundong Kong +2
Vision-Language Models (VLMs) have demonstrated strong performance across various multimodal tasks, where position encoding plays a vital role in modeling both the sequential struc…