6 papers
GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models
Zuyao You, Zhesong Yu, Mingyu Liu +3
In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamli…
Learning Accurate Segmentation Purely from Self-Supervision
Zuyao You, Zuxuan Wu, Yu-Gang Jiang
Accurately segmenting objects without any manual annotations remains one of the core challenges in computer vision. In this work, we introduce Selfment, a fully self-supervised fra…
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
Jiapeng Shi, Junke Wang, Zuyao You +2
This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and…
Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
Zuyao You, Zuxuan Wu
We present Seg-R1, a preliminary exploration of using reinforcement learning (RL) to enhance the pixel-level understanding and reasoning capabilities of large multimodal models (LM…
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
Zuyao You, Junke Wang, Lingyu Kong +2
We present Pix2Cap-COCO, the first panoptic pixel-level caption dataset designed to advance fine-grained visual understanding. To achieve this, we carefully design an automated ann…
FOCUS: Towards Universal Foreground Segmentation
Zuyao You, Lingyu Kong, Lingchen Meng +1
Foreground segmentation is a fundamental task in computer vision, encompassing various subdivision tasks. Previous research has typically designed task-specific architectures for e…