3 papers
cs.MM2026
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
Qihao Zhao, Yunqi Cao, Yangyu Huang +4
Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reas…
cs.CV2024
LTRL: Boosting Long-tail Recognition via Reflective Learning
Qihao Zhao, Yalun Dai, Shen Lin +3
In real-world scenarios, where knowledge distributions exhibit long-tail. Humans manage to master knowledge uniformly across imbalanced distributions, a feat attributed to their di…
cs.CV2024
LTGC: Long-tail Recognition via Leveraging LLMs-driven Generated Content
Qihao Zhao, Yalun Dai, Hao Li +3
Long-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper, w…