1 paper
Hyogun Lee, Soyeon Hong, Mujeen Sung +1
In this work, we tackle the problem of long-form video-language grounding (VLG). Given a long-form video and a natural language query, a model should temporally localize the precis…