activity
20192024
most citedDeepViT: Towards Deeper Vision Transformer

349 citations · 847 across the 9 of their papers we have counts for

collaborators

12 papers

cs.CV20224 cited

Diffusion Probabilistic Model Made Slim

Xingyi Yang, Daquan Zhou, Jiashi Feng +1

Despite the recent visually-pleasing results achieved, the massive computational cost has been a long-standing flaw for diffusion probabilistic models (DPMs), which, in turn, great…

cs.CV202214 cited

MagicMix: Semantic Mixing with Diffusion Models

Jun Hao Liew, Hanshu Yan, Daquan Zhou +1

Have you ever imagined what a corgi-alike coffee machine or a tiger-alike rabbit would look like? In this work, we attempt to answer these questions by exploring a new task called…

cs.CV202287 cited

MBEV: Multi-Camera Joint 3D Detection and Segmentation with Unified Birds-Eye View Representation

Enze Xie, Zhiding Yu, Daquan Zhou +5

In this paper, we propose MBEV, a unified framework that jointly performs 3D object detection and map segmentation in the Birds Eye View~(BEV) space with multi-camera image inp…

cs.CV202141 cited

Refiner: Refining Self-attention for Vision Transformers

Daquan Zhou, Yujun Shi, Bingyi Kang +6

Vision Transformers (ViTs) have shown competitive accuracy in image classification tasks compared with CNNs. Yet, they generally require much more data for model pre-training. Most…

cs.CV2021349 cited

DeepViT: Towards Deeper Vision Transformer

Daquan Zhou, Bingyi Kang, Xiaojie Jin +5

Vision transformers (ViTs) have been successfully applied in image classification tasks recently. In this paper, we show that, unlike convolution neural networks (CNNs)that can be…

cs.CV2021

All Tokens Matter: Token Labeling for Training Better Vision Transformers

Zihang Jiang, Qibin Hou, Li Yuan +5

In this paper, we present token labeling -- a new training objective for training high-performance vision transformers (ViTs). Different from the standard training objective of ViT…