activity
20162022
most citedMultiscale Vision Transformers

56 citations · 97 across the 6 of their papers we have counts for

collaborators

14 papers

cs.CV20221 cited

On the Importance of Asymmetry for Siamese Representation Learning

Xiao Wang, Haoqi Fan, Yuandong Tian +2

Many recent self-supervised frameworks for visual representation learning are based on certain forms of Siamese networks. Such networks are conceptually symmetric with two parallel…

cs.CV2021

A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning

Christoph Feichtenhofer, Haoqi Fan, Bo Xiong +2

We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simp…

cs.CV202156 cited

Multiscale Vision Transformers

Haoqi Fan, Bo Xiong, Karttikeya Mangalam +4

We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multisca…

cs.CV2021

Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories

Xitong Yang, Haoqi Fan, Lorenzo Torresani +2

The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We…

cs.CV2021

Multiview Pseudo-Labeling for Semi-supervised Learning from Video

Bo Xiong, Haoqi Fan, Kristen Grauman +1

We present a multiview pseudo-labeling approach to video learning, a novel framework that uses complementary views in the form of appearance and motion information for semi-supervi…

cs.CV2021

HiT: Hierarchical Transformer with Momentum Contrast for Video-Text Retrieval

Song Liu, Haoqi Fan, Shengsheng Qian +3

Video-Text Retrieval has been a hot research topic with the growth of multimedia data on the internet. Transformer for video-text learning has attracted increasing attention due to…