1 paper
Sangho Lee, Il Yong Chun, Hogun Park
Multi-modal transformers are rapidly gaining attention in video captioning tasks. Existing multi-modal video captioning methods typically extract a fixed number of frames, which ra…