3 papers
cs.CV2025
SAM2MOT: A Novel Paradigm of Multi-Object Tracking by Segmentation
Junjie Jiang, Zelin Wang, Manqi Zhao +2
Inspired by Segment Anything 2, which generalizes segmentation from images to videos, we propose SAM2MOT--a novel segmentation-driven paradigm for multi-object tracking that breaks…
cs.CV2025
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
Bingchao Wang, Zhiwei Ning, Jianyu Ding +5
CLIP has shown promising performance across many short-text tasks in a zero-shot manner. However, limited by the input length of the text encoder, CLIP struggles on under-stream ta…
cs.CV2025
MSF: Efficient Diffusion Model Via Multi-Scale Latent Factorize
Haohang Xu, Longyu Chen, Yichen Zhang +2
While diffusion-based generative models have made significant strides in visual content creation, conventional approaches face computational challenges, especially for high-resolut…