3 papers
cs.MM2026
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
Qihao Zhao, Yunqi Cao, Yangyu Huang +4
Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reas…
cs.LG2025
RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
Suhang Hu, Wei Hu, Yuhang Su +1
Vision-Language Models (VLMs) struggle with complex image annotation tasks, such as emotion classification and context-driven object detection, which demand sophisticated reasoning…
cs.CV2024
LTGC: Long-tail Recognition via Leveraging LLMs-driven Generated Content
Qihao Zhao, Yalun Dai, Hao Li +3
Long-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper, w…