2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2022
MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval
Yuying Ge, Yixiao Ge, Xihui Liu +5
Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast gl…
cs.CV2022★ 2 cited
Semantic-Aware Pretraining for Dense Video Captioning
Teng Wang, Zhu Liu, Feng Zheng +3
This report describes the details of our approach for the event dense-captioning task in ActivityNet Challenge 2021. We present a semantic-aware pretraining method for dense video…