1 paper
Houjing Wei, Yuting Shi, Naoya Inoue
Vision Large Language Models (VLLMs) usually take input as a concatenation of image token embeddings and text token embeddings and conduct causal modeling. However, their internal…