1 paper · 1 filter
Xuanyue Zhong, Yuqiang Xie, Guanqun Bi +2
Current video moment retrieval excels at action-centric tasks but struggles with narrative content. Models can see \textit{what is happening} but fail to reason \textit{why it matt…