1 paper
Nikos Papasarantopoulos, Shay B. Cohen
Research on text generation from multimodal inputs has largely focused on static images, and less on video data. In this paper, we propose a new task, narration generation, that is…