1 paper
Cheng Chen, Jingyu Zhou, Yifan Zhao +1
Understanding multi-label images remains a challenging task in computer vision. With the rapid progress of vision-language multimodal learning, vision-language models (VLMs) enable…