1 paper
Shaobo Ju, Haiyang Yu, Xuecheng Wu +6
Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by…