5 citations · 15 across the 18 of their papers we have counts for
1 paper · 1 filter
Kai Sun, Yushi Bai, Zhen Yang +4
Large Multimodal Models (LMMs) typically build on ViTs (e.g., CLIP), yet their training with simple random in-batch negatives limits the ability to capture fine-grained visual diff…