1 paper · 1 filter
Hyungyu Choi, Young Kyun Jang, Chanho Eom
Vision-language models such as CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions due to pre-t…