44 citations · 48 across the 2 of their papers we have counts for
1 paper · 1 filter
Shen Yan, Xuehan Xiong, Anurag Arnab +4
Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer…