3 citations · 6 across the 3 of their papers we have counts for
4 papers
SBAT: Video Captioning with Sparse Boundary-Aware Transformer
Tao Jin, Siyu Huang, Ming Chen +2
In this paper, we focus on the problem of applying the transformer structure to video captioning effectively. The vanilla transformer is proposed for uni-modal language generation…
Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning
Tao Jin, Siyu Huang, Yingming Li +1
This paper addresses the challenging task of video captioning which aims to generate descriptions for video data. Recently, the attention-based encoder-decoder structures have been…
Text Guided Person Image Synthesis
Xingran Zhou, Siyu Huang, Bin Li +3
This paper presents a novel method to manipulate the visual appearance (pose and attribute) of a person image according to natural language descriptions. Our method can be boiled d…
Tensor Decomposition via Variational Auto-Encoder
Bin Liu, Zenglin Xu, Yingming Li
Tensor decomposition is an important technique for capturing the high-order interactions among multiway data. Multi-linear tensor composition methods, such as the Tucker decomposit…