4 citations · 4 across the 1 of their papers we have counts for
3 papers
cs.CV2024★ 4 cited
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
Xingning Dong, Qingpei Guo, Tian Gan +5
We present a framework for learning cross-modal video representations by directly pre-training on raw data to facilitate various downstream video-text tasks. Our main contributions…
cs.CV2024
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
Xuzheng Yu, Chen Jiang, Wei Zhang +6
With the explosive growth of video data in real-world applications, a comprehensive representation of videos becomes increasingly important. In this paper, we address the problem o…
cs.CV2023
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
Chen Jiang, Hong Liu, Xuzheng Yu +8
In recent years, the explosion of web videos makes text-video retrieval increasingly essential and popular for video filtering, recommendation, and search. Text-video retrieval aim…