4 citations · 4 across the 1 of their papers we have counts for
1 paper
David Nukrai, Ron Mokady, Amir Globerson
We consider the task of image-captioning using only the CLIP model and additional text data at training time, and no additional captioned images. Our approach relies on the fact th…