214 citations · 776 across the 22 of their papers we have counts for
44 papers · 1 filter
Bootstrap Your Object Detector via Mixed Training
Mengde Xu, Zheng Zhang, Fangyun Wei +5
We introduce MixTraining, a new training paradigm for object detection that can improve the performance of existing detectors for free. MixTraining enhances data augmentation by ut…
Video Swin Transformer
Ze Liu, Jia Ning, Yue Cao +4
The vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchm…
Aligning Pretraining for Detection via Object-Level Contrastive Learning
Fangyun Wei, Yue Gao, Zhirong Wu +2
Image-level contrastive representation learning has proven to be highly effective as a generic model for transfer learning. Such generality for transfer learning, however, sacrific…
Neural Articulated Radiance Field
Atsuhiro Noguchi, Xiao Sun, Stephen Lin +1
We present Neural Articulated Radiance Field (NARF), a novel deformable 3D representation for articulated objects learned from images. While recent advances in 3D implicit represen…
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Ze Liu, Yutong Lin, Yue Cao +5
This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer fro…
Learning Monocular Depth in Dynamic Scenes via Instance-Aware Projection Consistency
Seokju Lee, Sunghoon Im, Stephen Lin +1
We present an end-to-end joint training framework that explicitly models 6-DoF motion of multiple dynamic objects, ego-motion and depth in a monocular camera setup without supervis…