1.6k citations · 2.1k across the 4 of their papers we have counts for
10 papers · 1 filter
Early Convolutions Help Transformers See Better
Tete Xiao, Mannat Singh, Eric Mintun +3
Vision transformer (ViT) models exhibit substandard optimizability. In particular, they are sensitive to the choice of optimizer (AdamW vs. SGD), optimizer hyperparameters, and tra…
Fast and Accurate Model Scaling
Piotr Dollár, Mannat Singh, Ross Girshick
In this work we analyze strategies for convolutional neural network scaling; that is, the process of scaling a base convolutional network to endow it with greater computational com…
Designing Network Design Spaces
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick +2
In this work, we present a new network design paradigm. Our goal is to help advance the understanding of network design and discover design principles that generalize across settin…
LVIS: A Dataset for Large Vocabulary Instance Segmentation
Agrim Gupta, Piotr Dollár, Ross Girshick
Progress on object detection is enabled by datasets that focus the research community's attention on open challenges. This process led us from simple images to complex scenes and f…
On Network Design Spaces for Visual Recognition
Ilija Radosavovic, Justin Johnson, Saining Xie +2
Over the past several years progress in designing better neural network architectures for visual recognition has been substantial. To help sustain this rate of progress, in this wo…
TensorMask: A Foundation for Dense Object Segmentation
Xinlei Chen, Ross Girshick, Kaiming He +1
Sliding-window object detectors that generate bounding-box object predictions over a dense, regular grid have advanced rapidly and proven popular. In contrast, modern instance segm…