1 paper
Yiqi Lin, Conghui He, Alex Jinpeng Wang +3
Despite CLIP being the foundation model in numerous vision-language applications, the CLIP suffers from a severe text spotting bias. Such bias causes CLIP models to `Parrot' the vi…