1 paper
Ji Ha Jang, Hayeon Kim, Chulwon Lee +2
CLIP (Contrastive Language-Image Pre-training) has become a de facto paradigm for image-text alignment, but it struggles with long-context descriptions (>77 tokens) due to absolute…