8 citations · 11 across the 3 of their papers we have counts for
3 papers
cs.CV2022
MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval
Yuying Ge, Yixiao Ge, Xihui Liu +5
Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast gl…
cs.CV2022★ 8 cited
UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection
Ye Liu, Siyuan Li, Yang Wu +3
Finding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era. Nev…
cs.CV2022★ 3 cited
All in One: Exploring Unified Video-Language Pre-training
Alex Jinpeng Wang, Yixiao Ge, Rui Yan +7
Mainstream Video-Language Pre-training models \cite{actbert,clipbert,violet} consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They purs…