1 paper
Zhengtao Zou, Ya Gao, Jiarui Guan +2
Large Vision-Language Models (LVLMs) typically process visual inputs as a prefix to the language decoder. As the model autoregressively generates text, this initial visual informat…