84 citations · 176 across the 10 of their papers we have counts for
9 papers · 1 filter
Exploring Motion and Appearance Information for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Pan Zhou +1
This paper addresses temporal sentence grounding. Previous works typically solve this task by learning frame-level video features and align them with the textual information. A maj…
HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot Learning
Shiming Chen, Guo-Sen Xie, Yang Liu +5
Zero-shot learning (ZSL) tackles the unseen class recognition problem, transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge…
Cross-Sentence Temporal and Semantic Relations in Video Activity Localisation
Jiabo Huang, Yang Liu, Shaogang Gong +1
Video activity localisation has recently attained increasing attention due to its practical values in automatically localising the most salient visual segments corresponding to the…
TEACHTEXT: CrossModal Generalized Distillation for Text-Video Retrieval
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu +4
In recent years, considerable progress on the task of text-video retrieval has been achieved by leveraging large-scale pretraining on visual and audio datasets to construct powerfu…
QuerYD: A video dataset with high-quality text and audio narrations
Andreea-Maria Oncescu, João F. Henriques, Yang Liu +2
We introduce QuerYD, a new large-scale dataset for retrieval and event localisation in video. A unique feature of our dataset is the availability of two audio tracks for each video…
The End-of-End-to-End: A Video Understanding Pentathlon Challenge (2020)
Samuel Albanie, Yang Liu, Arsha Nagrani +18
We present a new video understanding pentathlon challenge, an open competition held in conjunction with the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2020.…