Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
Yuan Yao, Qiushi Yang, Humen Zhong +5
Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large language models (MLLMs) exhibit st…
cs.CV2025
VA-AR: Learning Velocity-Aware Action Representations with Mixture of Window Attention
Jiangning Wei, Lixiong Qin, Bo Yu +7
Action recognition is a crucial task in artificial intelligence, with significant implications across various domains. We initially perform a comprehensive analysis of seven promin…
cs.CV2025
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
Lixiong Qin, Shilong Ou, Miaoxuan Zhang +6
Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable m…