67 citations · 71 across the 8 of their papers we have counts for
8 papers
Revisiting the Integration of Convolution and Attention for Vision Backbone
Lei Zhu, Xinjiang Wang, Wayne Zhang +1
Convolutions (Convs) and multi-head self-attentions (MHSAs) are typically considered alternatives to each other for building vision backbones. Although some works try to integrate…
Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
Lei Zhu, Fangyun Wei, Yanye Lu +1
In the realm of image quantization exemplified by VQGAN, the process encodes images into discrete tokens drawn from a codebook with a predefined size. Recent advancements, particul…
OpenSUN3D: 1st Workshop Challenge on Open-Vocabulary 3D Scene Understanding
Francis Engelmann, Ayca Takmaz, Jonas Schult +25
This report provides an overview of the challenge hosted at the OpenSUN3D Workshop on Open-Vocabulary 3D Scene Understanding held in conjunction with ICCV 2023. The goal of this wo…
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
Lei Zhu, Fangyun Wei, Yanye Lu
In this work, we investigate the potential of a large language model (LLM) to directly comprehend visual signals without the necessity of fine-tuning on multi-modal datasets. The f…
Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial Backpropagation
Yijun Yang, Angelica I. Aviles-Rivero, Huazhu Fu +3
Although convolutional neural networks (CNNs) have been proposed to remove adverse weather conditions in single images using a single set of pre-trained weights, they fail to resto…
Branches Mutual Promotion for End-to-End Weakly Supervised Semantic Segmentation
Lei Zhu, Hangzhou He, Xinliang Zhang +4
End-to-end weakly supervised semantic segmentation aims at optimizing a segmentation model in a single-stage training process based on only image annotations. Existing methods adop…