1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Yiqi Lin, Conghui He, Alex Jinpeng Wang +3
Despite CLIP being the foundation model in numerous vision-language applications, the CLIP suffers from a severe text spotting bias. Such bias causes CLIP models to `Parrot' the vi…