5 papers
HeartMuLa: A Family of Open Sourced Music Foundation Models
Dongchao Yang, Yuxin Xie, Yuguo Yin +26
We present a family of open-source Music Foundation Models designed to advance large-scale music understanding and generation across diverse tasks and modalities. Our framework con…
Joint Agent Memory and Exploration Learning via Novelty Signals
Shizuo Tian, Xiaohong Weng, Rui Kong +9
In open-ended environments, exploration is fundamental for autonomous agents, yet current language model agents struggle with this. Effective exploration requires memory, but retai…
TennisExpert: Towards Expert-Level Analytical Sports Video Understanding
Zhaoyu Liu, Xi Weng, Lianyu Hu +4
Tennis is one of the most widely followed sports, generating extensive broadcast footage with strong potential for professional analysis, automated coaching, and real-time commenta…
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
Xingjian Zhang, Xi Weng, Yihao Yue +3
Video behavior recognition and scene understanding are fundamental tasks in multimodal intelligence, serving as critical building blocks for numerous real-world applications. Throu…
Clustering Properties of Self-Supervised Learning
Xi Weng, Jianing An, Xudong Ma +5
Self-supervised learning (SSL) methods via joint embedding architectures have proven remarkably effective at capturing semantically rich representations with strong clustering prop…