1 paper
Wenwen Yu, Yuliang Liu, Xingkui Zhu +3
We exploit the potential of the large-scale Contrastive Language-Image Pretraining (CLIP) model to enhance scene text detection and spotting tasks, transforming it into a robust ba…