1 paper
Dahye Kim, Jungin Park, Jiyoung Lee +2
Given an untrimmed video and a language query depicting a specific temporal moment in the video, video grounding aims to localize the time interval by understanding the text and vi…