335 citations · 506 across the 13 of their papers we have counts for
1 paper · 1 filter
Wenyan Li, Raphael Tang, Chengzu Li +3
Vision--language models (VLMs) often process visual inputs through a pretrained vision encoder, followed by a projection into the language model's embedding space via a connector c…