2 papers
cs.CV2025
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
Baoyao Yang, Wanyun Li, Dixin Chen +3
This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises…
cs.CV2025
Expertized Caption Auto-Enhancement for Video-Text Retrieval
Baoyao Yang, Junxiang Chen, Wanyun Li +2
Video-text retrieval has been stuck in the information mismatch caused by personalized and inadequate textual descriptions of videos. The substantial information gap between the tw…