147 citations · 399 across the 16 of their papers we have counts for
Showing 2022 · cs.CVShow all
3 papers · 2 filters
cs.CV2022★ 1 cited
What Makes for Good Tokenizers in Vision Transformer?
Shengju Qian, Yi Zhu, Wenbo Li +2
The architecture of transformers, which recently witness booming applications in vision tasks, has pivoted against the widespread convolutional paradigm. Relying on the tokenizatio…
cs.CV2022★ 2 cited
MixGen: A New Multi-Modal Data Augmentation
Xiaoshuai Hao, Yi Zhu, Srikar Appalaraju +4
Data augmentation is a necessity to enhance data efficiency in deep learning. For vision-language pre-training, data is only augmented either for images or for text in previous wor…
cs.CV2022★ 1 cited
BigDetection: A Large-scale Benchmark for Improved Object Detector Pre-training
Likun Cai, Zhi Zhang, Yi Zhu +3
Multiple datasets and open challenges for object detection have been introduced in recent years. To build more general and powerful object detection systems, in this paper, we cons…