3 papers
cs.CV2019
Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network
Bairui Wang, Lin Ma, Wei Zhang +3
In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We const…
cs.CV2019
CAD-Net: A Context-Aware Detection Network for Objects in Remote Sensing Imagery
Gongjie Zhang, Shijian Lu, Wei Zhang
Accurate and robust detection of multi-class objects in optical remote sensing images is essential to many real-world applications such as urban planning, traffic control, searchin…
cs.CV2018
Reconstruction Network for Video Captioning
Bairui Wang, Lin Ma, Wei Zhang +1
In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of…