activity
20192023
most citedPyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

42 citations · 153 across the 20 of their papers we have counts for

collaborators
Showing 2022 · cs.CVShow all

8 papers · 2 filters

cs.CV2022★ 5 cited

LAB-Net: LAB Color-Space Oriented Lightweight Network for Shadow Removal

Hong Yang, Gongrui Nan, Mingbao Lin +4

This paper focuses on the limitations of current over-parameterized shadow removal models. We present a novel lightweight deep neural network that processes shadow images in the LA…

cs.CV2022

Efficient Decoder-free Object Detection with Transformers

Peixian Chen, Mengdan Zhang, Yunhang Shen +5

Vision transformers (ViTs) are changing the landscape of object detection approaches. A natural usage of ViTs in detection is to replace the CNN-based backbone with a transformer-b…

cs.CV2022★ 7 cited

Open Vocabulary Object Detection with Proposal Mining and Prediction Equalization

Peixian Chen, Kekai Sheng, Mengdan Zhang +5

Open-vocabulary object detection (OVD) aims to scale up vocabulary size to detect objects of novel categories beyond the training vocabulary. Recent work resorts to the rich knowle…

cs.CV2022★ 42 cited

PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

Yuting Gao, Jinfeng Liu, Zihan Xu +4

Large-scale vision-language pre-training has achieved promising results on downstream tasks. Existing methods highly rely on the assumption that the image-text pairs crawled from t…

cs.CV2022

Super Vision Transformer

Mingbao Lin, Mengzhao Chen, Yuxin Zhang +3

We attempt to reduce the computational costs in vision transformers (ViTs), which increase quadratically in the token number. We present a novel training paradigm that trains only…

cs.CV2022

Training-free Transformer Architecture Search

Qinqin Zhou, Kekai Sheng, Xiawu Zheng +5

Recently, Vision Transformer (ViT) has achieved remarkable success in several computer vision tasks. The progresses are highly relevant to the architecture design, then it is worth…