14 citations · 16 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 14 cited
Vision-Language Pre-Training with Triple Contrastive Learning
Jinyu Yang, Jiali Duan, Son Tran +6
Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attrib…
cs.CV2021★ 2 cited
Improving Apparel Detection with Category Grouping and Multi-grained Branches
Qing Tian, Sampath Chanda, K C Amit Kumar +1
Training an accurate object detector is expensive and time-consuming. One main reason lies in the laborious labeling process, i.e., annotating category and bounding box information…