16 citations · 16 across the 5 of their papers we have counts for
1 paper · 1 filter
Moon Ye-Bin, Nam Hyeon-Woo, Wonseok Choi +1
Vision language models (VLMs) perceive the world through a combination of a visual encoder and a large language model (LLM). The visual encoder, pre-trained on large-scale vision-t…