1 paper
Ruchen Liu, Yi Yang, Yiming Xu +3
LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show tha…