1 paper
Sohee Kim, Jisu Kang, Dunam Kim +1
In this paper, we demonstrate that CLIP can also be adapted to downstream tasks where its vision-language alignment is suboptimally learned during pre-training on web-crawled data,…