8 citations · 8 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 8 cited
MMViT: Multiscale Multiview Vision Transformers
Yuchen Liu, Natasha Ong, Kaiyan Peng +8
We present Multiscale Multiview Vision Transformers (MMViT), which introduces multiscale feature maps and multiview encodings to transformer models. Our model encodes different vie…
cs.CV2023
SVT: Supertoken Video Transformer for Efficient Video Understanding
Chenbin Pan, Rui Hou, Hanchao Yu +3
Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video conte…