353 citations · 759 across the 8 of their papers we have counts for
28 papers
Early Convolutions Help Transformers See Better
Tete Xiao, Mannat Singh, Eric Mintun +3
Vision transformer (ViT) models exhibit substandard optimizability. In particular, they are sensitive to the choice of optimizer (AdamW vs. SGD), optimizer hyperparameters, and tra…
A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong +2
We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simp…
Boundary IoU: Improving Object-Centric Image Segmentation Evaluation
Bowen Cheng, Ross Girshick, Piotr Dollár +2
We present Boundary IoU (Intersection-over-Union), a new segmentation evaluation measure focused on boundary quality. We perform an extensive analysis across different error types…
Fast and Accurate Model Scaling
Piotr Dollár, Mannat Singh, Ross Girshick
In this work we analyze strategies for convolutional neural network scaling; that is, the process of scaling a base convolutional network to endow it with greater computational com…
Large scale weakly and semi-supervised learning for low-resource video ASR
Kritika Singh, Vimal Manohar, Alex Xiao +7
Many semi- and weakly-supervised approaches have been investigated for overcoming the labeling cost of building high quality speech recognition systems. On the challenging task of…
Designing Network Design Spaces
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick +2
In this work, we present a new network design paradigm. Our goal is to help advance the understanding of network design and discover design principles that generalize across settin…