17 citations · 17 across the 2 of their papers we have counts for
3 papers
cs.CV2022
Visual Commonsense-aware Representation Network for Video Captioning
Pengpeng Zeng, Haonan Zhang, Lianli Gao +3
Generating consecutive descriptions for videos, i.e., Video Captioning, requires taking full advantage of visual representation along with the generation process. Existing video ca…
cs.CV2018★ 17 cited
Hierarchical LSTMs with Adaptive Attention for Visual Captioning
Jingkuan Song, Xiangpeng Li, Lianli Gao +1
Recent progress has been made in using attention based encoder-decoder framework for image and video captioning. Most existing decoders apply the attention mechanism to every gener…
cs.CV2018
Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder
Jingkuan Song, Hanwang Zhang, Xiangpeng Li +3
Existing video hash functions are built on three isolated stages: frame pooling, relaxed learning, and binarization, which have not adequately explored the temporal order of video…