17 citations · 17 across the 2 of their papers we have counts for
3 papers · 1 filter
Visual Commonsense-aware Representation Network for Video Captioning
Pengpeng Zeng, Haonan Zhang, Lianli Gao +3
Generating consecutive descriptions for videos, i.e., Video Captioning, requires taking full advantage of visual representation along with the generation process. Existing video ca…
Hierarchical LSTMs with Adaptive Attention for Visual Captioning
Jingkuan Song, Xiangpeng Li, Lianli Gao +1
Recent progress has been made in using attention based encoder-decoder framework for image and video captioning. Most existing decoders apply the attention mechanism to every gener…
Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder
Jingkuan Song, Hanwang Zhang, Xiangpeng Li +3
Existing video hash functions are built on three isolated stages: frame pooling, relaxed learning, and binarization, which have not adequately explored the temporal order of video…