1 paper
Oscar Chew, Hsiao-Ying Huang, Kunal Jain +3
Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to a…