5 citations · 9 across the 5 of their papers we have counts for
1 paper · 1 filter
Mengmeng Xu, Erhan Gundogdu, Maksim Lapin +3
Long-form video understanding requires designing approaches that are able to temporally localize activities or language. End-to-end training for such tasks is limited by the comput…