2 citations · 2 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
SVAC: Scaling Is All You Need For Referring Video Object Segmentation
Li Zhang, Haoxiang Gao, Zhihao Zhang +2
Referring Video Object Segmentation (RVOS) aims to segment target objects in video sequences based on natural language descriptions. While recent advances in Multi-modal Large Lang…
cs.CV2025
OmniCam: Unified Multimodal Video Generation via Camera Control
Xiaoda Yang, Jiayang Xu, Kaixuan Luan +9
Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as co…
cs.CV2025
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
Haoxiang Gao, Li Zhang, Yu Zhao +2
Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand com…