1 paper
Yapeng Tian, Chenxiao Guan, Justin Goodman +2
Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory a…