5 papers
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
Qihao Zhao, Yunqi Cao, Yangyu Huang +4
Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reas…
LTRL: Boosting Long-tail Recognition via Reflective Learning
Qihao Zhao, Yalun Dai, Shen Lin +3
In real-world scenarios, where knowledge distributions exhibit long-tail. Humans manage to master knowledge uniformly across imbalanced distributions, a feat attributed to their di…
LTGC: Long-tail Recognition via Leveraging LLMs-driven Generated Content
Qihao Zhao, Yalun Dai, Hao Li +3
Long-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper, w…
A discussion on numerical shock stability of unstructured finite volume method: Riemann solvers and limiters
Fan Zhang, Zhichao Yuan, Jun Liu
Numerical shock instability is a complexity which may occur in supersonic simulations. Riemann solver is usually the crucial factor that affects both the computation accuracy and n…
MDCS: More Diverse Experts with Consistency Self-distillation for Long-tailed Recognition
Qihao Zhao, Chen Jiang, Wei Hu +2
Recently, multi-expert methods have led to significant improvements in long-tail recognition (LTR). We summarize two aspects that need further enhancement to contribute to LTR boos…