8 citations · 9 across the 2 of their papers we have counts for
1 paper · 1 filter
Yuchen Liu, Natasha Ong, Kaiyan Peng +8
We present Multiscale Multiview Vision Transformers (MMViT), which introduces multiscale feature maps and multiview encodings to transformer models. Our model encodes different vie…