1 paper
Donghee Lee, Rui Cai, Zhe Zhao
Large vision-language models (LVLMs) are typically trained using autoregressive language modeling objectives, which align visual representations with linguistic space. While effect…