1 paper
Xingyi Zhou, Anurag Arnab, Chen Sun +1
We propose a new task and model for dense video object captioning -- detecting, tracking and captioning trajectories of objects in a video. This task unifies spatial and temporal l…