14 citations · 14 across the 4 of their papers we have counts for
1 paper · 1 filter
Wilson Yan, Volodymyr Mnih, Aleksandra Faust +3
Efficient video tokenization remains a key bottleneck in learning general purpose vision models that are capable of processing long video sequences. Prevailing approaches are restr…