56 citations · 97 across the 6 of their papers we have counts for
14 papers
On the Importance of Asymmetry for Siamese Representation Learning
Xiao Wang, Haoqi Fan, Yuandong Tian +2
Many recent self-supervised frameworks for visual representation learning are based on certain forms of Siamese networks. Such networks are conceptually symmetric with two parallel…
A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong +2
We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simp…
Multiscale Vision Transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam +4
We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multisca…
Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories
Xitong Yang, Haoqi Fan, Lorenzo Torresani +2
The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We…
Multiview Pseudo-Labeling for Semi-supervised Learning from Video
Bo Xiong, Haoqi Fan, Kristen Grauman +1
We present a multiview pseudo-labeling approach to video learning, a novel framework that uses complementary views in the form of appearance and motion information for semi-supervi…
HiT: Hierarchical Transformer with Momentum Contrast for Video-Text Retrieval
Song Liu, Haoqi Fan, Shengsheng Qian +3
Video-Text Retrieval has been a hot research topic with the growth of multimedia data on the internet. Transformer for video-text learning has attracted increasing attention due to…