2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Tianyu Chen, Xingcheng Fu, Yisen Gao +5
Modern vision-language models (VLMs) develop patch embedding and convolution backbone within vector space, especially Euclidean ones, at the very founding. When expanding VLMs to a…