42 citations · 153 across the 20 of their papers we have counts for
8 papers · 2 filters
LAB-Net: LAB Color-Space Oriented Lightweight Network for Shadow Removal
Hong Yang, Gongrui Nan, Mingbao Lin +4
This paper focuses on the limitations of current over-parameterized shadow removal models. We present a novel lightweight deep neural network that processes shadow images in the LA…
Efficient Decoder-free Object Detection with Transformers
Peixian Chen, Mengdan Zhang, Yunhang Shen +5
Vision transformers (ViTs) are changing the landscape of object detection approaches. A natural usage of ViTs in detection is to replace the CNN-based backbone with a transformer-b…
Open Vocabulary Object Detection with Proposal Mining and Prediction Equalization
Peixian Chen, Kekai Sheng, Mengdan Zhang +5
Open-vocabulary object detection (OVD) aims to scale up vocabulary size to detect objects of novel categories beyond the training vocabulary. Recent work resorts to the rich knowle…
PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining
Yuting Gao, Jinfeng Liu, Zihan Xu +4
Large-scale vision-language pre-training has achieved promising results on downstream tasks. Existing methods highly rely on the assumption that the image-text pairs crawled from t…
Super Vision Transformer
Mingbao Lin, Mengzhao Chen, Yuxin Zhang +3
We attempt to reduce the computational costs in vision transformers (ViTs), which increase quadratically in the token number. We present a novel training paradigm that trains only…
Training-free Transformer Architecture Search
Qinqin Zhou, Kekai Sheng, Xiawu Zheng +5
Recently, Vision Transformer (ViT) has achieved remarkable success in several computer vision tasks. The progresses are highly relevant to the architecture design, then it is worth…