1 paper
Suhang Hu, Wei Hu, Yuhang Su +1
Vision-Language Models (VLMs) struggle with complex image annotation tasks, such as emotion classification and context-driven object detection, which demand sophisticated reasoning…