11 citations · 16 across the 2 of their papers we have counts for
2 papers
cs.CV2020★ 5 cited
A Simple Yet Effective Method for Video Temporal Grounding with Cross-Modality Attention
Binjie Zhang, Yu Li, Chun Yuan +3
The task of language-guided video temporal grounding is to localize the particular video clip corresponding to a query sentence in an untrimmed video. Though progress has been made…
cs.CV2019★ 11 cited
ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Zhou Yu, Dejing Xu, Jun Yu +4
Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to…