1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Raghav Goyal, Effrosyni Mavroudi, Xitong Yang +5
Video understanding tasks take many forms, from action detection to visual query localization and spatio-temporal grounding of sentences. These tasks differ in the type of inputs (…