3 papers
cs.CV2026
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation
Yanhui Chen, Baoyao Yang, Siqi Liu +1
SAM3 advances open-vocabulary semantic segmentation by introducing a prompt-driven mask generation paradigm. However, in multi-class open-vocabulary scenarios, masks generated inde…
cs.CV2025
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
Baoyao Yang, Wanyun Li, Dixin Chen +3
This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises…
cs.CV2025
Expertized Caption Auto-Enhancement for Video-Text Retrieval
Baoyao Yang, Junxiang Chen, Wanyun Li +2
Video-text retrieval has been stuck in the information mismatch caused by personalized and inadequate textual descriptions of videos. The substantial information gap between the tw…