1 paper
Zelin Zheng, Xinyan Liu, Ruixin Li +4
Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often leading to brittle num…