activity
20152021
most citedEarly Convolutions Help Transformers See Better

353 citations · 759 across the 8 of their papers we have counts for

collaborators

28 papers

cs.CV2021353 cited

Early Convolutions Help Transformers See Better

Tete Xiao, Mannat Singh, Eric Mintun +3

Vision transformer (ViT) models exhibit substandard optimizability. In particular, they are sensitive to the choice of optimizer (AdamW vs. SGD), optimizer hyperparameters, and tra…

cs.CV2021

A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning

Christoph Feichtenhofer, Haoqi Fan, Bo Xiong +2

We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simp…

cs.CV202128 cited

Boundary IoU: Improving Object-Centric Image Segmentation Evaluation

Bowen Cheng, Ross Girshick, Piotr Dollár +2

We present Boundary IoU (Intersection-over-Union), a new segmentation evaluation measure focused on boundary quality. We perform an extensive analysis across different error types…

cs.CV2021

Fast and Accurate Model Scaling

Piotr Dollár, Mannat Singh, Ross Girshick

In this work we analyze strategies for convolutional neural network scaling; that is, the process of scaling a base convolutional network to endow it with greater computational com…

eess.AS2020

Large scale weakly and semi-supervised learning for low-resource video ASR

Kritika Singh, Vimal Manohar, Alex Xiao +7

Many semi- and weakly-supervised approaches have been investigated for overcoming the labeling cost of building high quality speech recognition systems. On the challenging task of…

cs.CV2020

Designing Network Design Spaces

Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick +2

In this work, we present a new network design paradigm. Our goal is to help advance the understanding of network design and discover design principles that generalize across settin…