1 paper · 1 filter
Jiayu Li, Rajesh Gangireddy, Samet Akcay +2
Vision-Language Models (VLMs) learn powerful multimodal representations through large-scale image-text pretraining, but adapting them to hierarchical classification is underexplore…