1 paper · 1 filter
Ethan Baron, Idan Tankel, Peter Tu +1
In this study, we define and tackle zero shot "real" classification by description, a novel task that evaluates the ability of Vision-Language Models (VLMs) like CLIP to classify o…