4 papers
Grounded Video Description
Luowei Zhou, Yannis Kalantidis, Xinlei Chen +2
Video description is one of the most challenging problems in vision and language understanding due to the large variability both on the video and language side. Models, hence, typi…
Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition
Hao Huang, Luowei Zhou, Wei Zhang +2
Video action recognition, a critical problem in video understanding, has been gaining increasing attention. To identify actions induced by complex object-object interactions, we ne…
Weakly-Supervised Video Object Grounding from Text by Loss Weighting and Object Interaction
Luowei Zhou, Nathan Louis, Jason J. Corso
We study weakly-supervised video object grounding: given a video segment and a corresponding descriptive sentence, the goal is to localize objects that are mentioned from the sente…
End-to-End Dense Video Captioning with Masked Transformer
Luowei Zhou, Yingbo Zhou, Jason J. Corso +2
Dense video captioning aims to generate text descriptions for all events in an untrimmed video. This involves both detecting and describing events. Therefore, all previous methods…