2 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Tejas Srinivasan, Xiang Ren, Jesse Thomason
Aligning image and text encoders from scratch using contrastive learning requires large amounts of paired image-text data. We alleviate this need by aligning individually pre-train…