2 citations · 2 across the 1 of their papers we have counts for
1 paper
Yanan Zhang, Jiangmeng Li, Lixiang Liu +1
Foundational Vision-Language models such as CLIP have exhibited impressive generalization in downstream tasks. However, CLIP suffers from a two-level misalignment issue, i.e., task…