activity
20242026
collaborators

8 papers

cs.CV2026

AnyTrack: Unifying Visual Object Tracking with Any Modalities

Hao Li, Yunzhi Zhuge, Wenning Hao +4

Visual object tracking aims to continuously locate specific targets within sequential frames, evolving from single-modal methods to multi-modal ones. However, existing multi-modal…

cs.CV2026

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs

Yuhao Wang, Mu Qiao, Haiwen Diao +5

Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token sequences. Training-free visual token reduction provides an efficient solution.…

cs.CV2025

Parameter Aware Mamba Model for Multi-task Dense Prediction

Xinzhuo Yu, Yunzhi Zhuge, Sitong Gong +3

Understanding the inter-relations and interactions between tasks is crucial for multi-task dense prediction. Existing methods predominantly utilize convolutional layers and attenti…

cs.CV2025

Complementary and Contrastive Learning for Audio-Visual Segmentation

Sitong Gong, Yunzhi Zhuge, Lu Zhang +2

Audio-Visual Segmentation (AVS) aims to generate pixel-wise segmentation maps that correlate with the auditory signals of objects. This field has seen significant progress with num…

cs.CV2025

Reinforcing Video Reasoning Segmentation to Think Before It Segments

Sitong Gong, Lu Zhang, Yunzhi Zhuge +3

Video reasoning segmentation (VRS) endeavors to delineate referred objects in videos guided by implicit instructions that encapsulate human intent and temporal logic. Previous appr…

cs.CV2025

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

Sitong Gong, Yunzhi Zhuge, Lu Zhang +3

Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial…