1 paper
Moon Ye-Bin, Nam Hyeon-Woo, Wonseok Choi +1
Vision language models (VLMs) perceive the world through a combination of a visual encoder and a large language model (LLM). The visual encoder, pre-trained on large-scale vision-t…