activity
20172025
most citedToken Merging: Your ViT But Faster

65 citations · 254 across the 29 of their papers we have counts for

collaborators
Showing 2022Show all

7 papers · 1 filter

cs.CV2022★ 4 cited

3D-Aware Encoding for Style-based Neural Radiance Fields

Yu-Jhe Li, Tao Xu, Bichen Wu +6

We tackle the task of NeRF inversion for style-based neural radiance fields, (e.g., StyleNeRF). In the task, we aim to learn an inversion function to project an input image to the…

cs.CV2022★ 3 cited

Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference

Haoran You, Yunyang Xiong, Xiaoliang Dai +5

Vision Transformers (ViTs) have shown impressive performance but still require a high computation cost as compared to convolutional neural networks (CNNs), one reason is that ViTs'…

cs.CV2022★ 65 cited

Token Merging: Your ViT But Faster

Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai +3

We introduce Token Merging (ToMe), a simple method to increase the throughput of existing ViT models without needing to train. ToMe gradually combines similar tokens in a transform…

cs.CV2022★ 10 cited

Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP

Feng Liang, Bichen Wu, Xiaoliang Dai +6

Open-vocabulary semantic segmentation aims to segment an image into semantic regions according to text descriptions, which may not have been seen during training. Recent two-stage…

cs.CV2022★ 2 cited

Hydra Attention: Efficient Attention with Many Heads

Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai +2

While transformers have begun to dominate many tasks in vision, applying them to large images is still computationally difficult. A large reason for this is that self-attention sca…

cs.CV2022

Open-Set Semi-Supervised Object Detection

Yen-Cheng Liu, Chih-Yao Ma, Xiaoliang Dai +4

Recent developments for Semi-Supervised Object Detection (SSOD) have shown the promise of leveraging unlabeled data to improve an object detector. However, thus far these methods h…