activity
20222024
most citedBiFormer: Vision Transformer with Bi-Level Routing Attention

67 citations · 71 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2024

Revisiting the Integration of Convolution and Attention for Vision Backbone

Lei Zhu, Xinjiang Wang, Wayne Zhang +1

Convolutions (Convs) and multi-head self-attentions (MHSAs) are typically considered alternatives to each other for building vision backbones. Although some works try to integrate…

cs.CV20241 cited

Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Lei Zhu, Fangyun Wei, Yanye Lu +1

In the realm of image quantization exemplified by VQGAN, the process encodes images into discrete tokens drawn from a codebook with a predefined size. Recent advancements, particul…

cs.CV20241 cited

OpenSUN3D: 1st Workshop Challenge on Open-Vocabulary 3D Scene Understanding

Francis Engelmann, Ayca Takmaz, Jonas Schult +25

This report provides an overview of the challenge hosted at the OpenSUN3D Workshop on Open-Vocabulary 3D Scene Understanding held in conjunction with ICCV 2023. The goal of this wo…

cs.CV2024

Beyond Text: Frozen Large Language Models in Visual Signal Comprehension

Lei Zhu, Fangyun Wei, Yanye Lu

In this work, we investigate the potential of a large language model (LLM) to directly comprehend visual signals without the necessity of fine-tuning on multi-modal datasets. The f…

cs.CV20231 cited

Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial Backpropagation

Yijun Yang, Angelica I. Aviles-Rivero, Huazhu Fu +3

Although convolutional neural networks (CNNs) have been proposed to remove adverse weather conditions in single images using a single set of pre-trained weights, they fail to resto…

cs.CV20231 cited

Branches Mutual Promotion for End-to-End Weakly Supervised Semantic Segmentation

Lei Zhu, Hangzhou He, Xinliang Zhang +4

End-to-end weakly supervised semantic segmentation aims at optimizing a segmentation model in a single-stage training process based on only image annotations. Existing methods adop…