4 citations · 4 across the 1 of their papers we have counts for
4 papers
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
Xingning Dong, Qingpei Guo, Tian Gan +5
We present a framework for learning cross-modal video representations by directly pre-training on raw data to facilitate various downstream video-text tasks. Our main contributions…
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
Xuzheng Yu, Chen Jiang, Wei Zhang +6
With the explosive growth of video data in real-world applications, a comprehensive representation of videos becomes increasingly important. In this paper, we address the problem o…
Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
Chen Jiang, Kaiming Huang, Sifeng He +9
With the explosive growth of web videos in recent years, large-scale Content-Based Video Retrieval (CBVR) becomes increasingly essential in video filtering, recommendation, and cop…
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
Chen Jiang, Hong Liu, Xuzheng Yu +8
In recent years, the explosion of web videos makes text-video retrieval increasingly essential and popular for video filtering, recommendation, and search. Text-video retrieval aim…