2 citations · 2 across the 1 of their papers we have counts for
1 paper
Xiaohu Jiang, Yixiao Ge, Yuying Ge +3
Image-text training like CLIP has dominated the pretraining of vision foundation models in recent years. Subsequent efforts have been made to introduce region-level visual learning…