activity
20182021
most citedMultiscale Vision Transformers

56 citations · 66 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2021

A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning

Christoph Feichtenhofer, Haoqi Fan, Bo Xiong +2

We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simp…

cs.CV202156 cited

Multiscale Vision Transformers

Haoqi Fan, Bo Xiong, Karttikeya Mangalam +4

We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multisca…

cs.CV20215 cited

Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos

Yanghao Li, Tushar Nagarajan, Bo Xiong +1

We introduce an approach for pre-training egocentric video models using large-scale third-person video datasets. Learning from purely egocentric data is limited by low dataset scal…

cs.CV2021

Multiview Pseudo-Labeling for Semi-supervised Learning from Video

Bo Xiong, Haoqi Fan, Kristen Grauman +1

We present a multiview pseudo-labeling approach to video learning, a novel framework that uses complementary views in the form of appearance and motion information for semi-supervi…

cs.CV20195 cited

Less is More: Learning Highlight Detection from Video Duration

Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram +1

Highlight detection has the potential to significantly ease video browsing, but existing methods often suffer from expensive supervision requirements, where human viewers must manu…

cs.CV2018

Pixel Objectness: Learning to Segment Generic Objects Automatically in Images and Videos

Bo Xiong, Suyog Dutt Jain, Kristen Grauman

We propose an end-to-end learning framework for segmenting generic objects in both images and videos. Given a novel image or video, our approach produces a pixel-level mask for all…