1 paper · 1 filter
Yuhui Zeng, Xinyu Mao, Xiaokun Liu +4
Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this inter…