1 paper
William Yicheng Zhu, Keren Ye, Junjie Ke +4
Recognizing and disentangling visual attributes from objects is a foundation to many computer vision applications. While large vision language representations like CLIP had largely…