1 paper
MinJu Jeon, Si-Woo Kim, Ye-Chan Kim +2
Dense video captioning aims to temporally localize events in video and generate captions for each event. While recent works propose end-to-end models, they suffer from two limitati…