activity
20212023
most citedMViTv2: Improved Multiscale Vision Transformers for Classification and Detection

29 citations · 47 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2023★ 8 cited

MMViT: Multiscale Multiview Vision Transformers

Yuchen Liu, Natasha Ong, Kaiyan Peng +8

We present Multiscale Multiview Vision Transformers (MMViT), which introduces multiscale feature maps and multiview encodings to transformer models. Our model encodes different vie…

cs.CV2023

Reversible Vision Transformers

Karttikeya Mangalam, Haoqi Fan, Yanghao Li +4

We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By decoupling the GPU memory requirement from the depth of the model, Reve…

cs.IR2022

Normalized Contrastive Learning for Text-Video Retrieval

Yookoon Park, Mahmoud Azab, Bo Xiong +4

Cross-modal contrastive learning has led the recent advances in multimodal retrieval with its simplicity and effectiveness. In this work, however, we reveal that cross-modal contra…

cs.CV2022★ 9 cited

MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition

Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam +4

While today's video recognition systems parse snapshots or short clips accurately, they cannot connect the dots and reason across a longer range of time yet. Most existing video ar…

cs.CV2021★ 29 cited

MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Yanghao Li, Chao-Yuan Wu, Haoqi Fan +4

In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection. We present an improved ve…

cs.CV2021★ 1 cited

PyTorchVideo: A Deep Learning Library for Video Understanding

Haoqi Fan, Tullie Murrell, Heng Wang +13

We introduce PyTorchVideo, an open-source deep-learning library that provides a rich set of modular, efficient, and reproducible components for a variety of video understanding tas…