1 paper
Hyun Lee, Hyemin Jeong, Yejin Kim +4
Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding…