1 paper
Lingfeng Qiao, Chen Wu, Ye Liu +3
Multimodal headline utilizes both video frames and transcripts to generate the natural language title of the videos. Due to a lack of large-scale, manually annotated data, the task…