1 paper · 1 filter
Chau Truong, Hieu Ta Quang, Dung D. Le
Vision-language models like CLIP show impressive ability to align images and text, but their training on short, concise captions makes them struggle with lengthy, detailed descript…