1 paper
Xiaoxing Hu, Kaicheng Yang, Ziyang Gong +6
The original CLIP text encoder is limited by a maximum input length of 77 tokens, which hampers its ability to effectively process long texts and perform fine-grained semantic unde…