15 citations · 18 across the 2 of their papers we have counts for
1 paper · 1 filter
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo +5
In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architec…