5 citations · 5 across the 1 of their papers we have counts for
1 paper
Chih-Yao Ma, Asim Kadav, Iain Melvin +3
We address the problem of video captioning by grounding language generation on object interactions in the video. Existing work mostly focuses on overall scene understanding with of…