1 paper
Chau Truong, Hieu Ta Quang, Dung D. Le
Vision-language models like CLIP show impressive ability to align images and text, but their training on short, concise captions makes them struggle with lengthy, detailed descript…