56 citations · 66 across the 4 of their papers we have counts for
7 papers
A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong +2
We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simp…
Multiscale Vision Transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam +4
We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multisca…
Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos
Yanghao Li, Tushar Nagarajan, Bo Xiong +1
We introduce an approach for pre-training egocentric video models using large-scale third-person video datasets. Learning from purely egocentric data is limited by low dataset scal…
Multiview Pseudo-Labeling for Semi-supervised Learning from Video
Bo Xiong, Haoqi Fan, Kristen Grauman +1
We present a multiview pseudo-labeling approach to video learning, a novel framework that uses complementary views in the form of appearance and motion information for semi-supervi…
Less is More: Learning Highlight Detection from Video Duration
Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram +1
Highlight detection has the potential to significantly ease video browsing, but existing methods often suffer from expensive supervision requirements, where human viewers must manu…
Pixel Objectness: Learning to Segment Generic Objects Automatically in Images and Videos
Bo Xiong, Suyog Dutt Jain, Kristen Grauman
We propose an end-to-end learning framework for segmenting generic objects in both images and videos. Given a novel image or video, our approach produces a pixel-level mask for all…