2 papers
cs.CV2025
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
Ye Sun, Hao Zhang, Henghui Ding +3
Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires…
cs.CV2024
UnSeg: One Universal Unlearnable Example Generator is Enough against All Image Segmentation
Ye Sun, Hao Zhang, Tiehua Zhang +2
Image segmentation is a crucial vision task that groups pixels within an image into semantically meaningful segments, which is pivotal in obtaining a fine-grained understanding of…