1 paper
Zhilei Shu, Shangwen Zhu, Zihang Liang +10
Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and ove…