activity
20182022
most citedFine-grained Video-Text Retrieval with Hierarchical Graph Reasoning

26 citations · 48 across the 6 of their papers we have counts for

collaborators

9 papers

cs.CV2022

Progressive Learning for Image Retrieval with Hybrid-Modality Queries

Yida Zhao, Yuqing Song, Qin Jin

Image retrieval with hybrid-modality queries, also known as composing text and image for image retrieval (CTI-IR), is a retrieval task where the search intention is expressed in a…

cs.CV2021

WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training

Yuqi Huo, Manli Zhang, Guangzhen Liu +32

Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…

cs.CV202010 cited

The End-of-End-to-End: A Video Understanding Pentathlon Challenge (2020)

Samuel Albanie, Yang Liu, Arsha Nagrani +18

We present a new video understanding pentathlon challenge, an open competition held in conjunction with the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2020.…

cs.CV20203 cited

Team RUC_AIM3 Technical Report at Activitynet 2020 Task 2: Exploring Sequential Events Detection for Dense Video Captioning

Yuqing Song, Shizhe Chen, Yida Zhao +1

Detecting meaningful events in an untrimmed video is essential for dense video captioning. In this work, we propose a novel and simple model for event sequence generation and explo…

cs.CV202026 cited

Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning

Shizhe Chen, Yida Zhao, Qin Jin +1

Cross-modal retrieval between videos and texts has attracted growing attentions due to the rapid emergence of videos on the web. The current dominant approach for this problem is t…

cs.CV2019

Integrating Temporal and Spatial Attentions for VATEX Video Captioning Challenge 2019

Shizhe Chen, Yida Zhao, Yuqing Song +2

This notebook paper presents our model in the VATEX video captioning challenge. In order to capture multi-level aspects in the video, we propose to integrate both temporal and spat…