3 citations · 3 across the 1 of their papers we have counts for
1 paper
Tao Jin, Siyu Huang, Ming Chen +2
In this paper, we focus on the problem of applying the transformer structure to video captioning effectively. The vanilla transformer is proposed for uni-modal language generation…