activity
20172023
most citedBiSeNet V2: Bilateral Network with Guided Aggregation for Real-time Semantic Segmentation

115 citations · 733 across the 55 of their papers we have counts for

collaborators

89 papers

cs.CV20235 cited

Learning Conditional Attributes for Compositional Zero-Shot Learning

Qingsheng Wang, Lingqiao Liu, Chenchen Jing +4

Compositional Zero-Shot Learning (CZSL) aims to train models to recognize novel compositional concepts based on learned concepts such as attribute-object combinations. One of the c…

cs.CV2022

FoPro: Few-Shot Guided Robust Webly-Supervised Prototypical Learning

Yulei Qin, Xingyu Chen, Chao Chen +5

Recently, webly supervised learning (WSL) has been studied to leverage numerous and accessible data from the Internet. Most existing methods focus on learning noise-robust models f…

cs.CV20221 cited

Learning from partially labeled data for multi-organ and tumor segmentation

Yutong Xie, Jianpeng Zhang, Yong Xia +1

Medical image benchmarks for the segmentation of organs and tumors suffer from the partially labeling issue due to its intensive cost of labor and expertise. Current mainstream app…

cs.CV20228 cited

Hierarchical Normalization for Robust Monocular Depth Estimation

Chi Zhang, Wei Yin, Zhibin Wang +3

In this paper, we address monocular depth estimation with deep neural networks. To enable training of deep monocular estimation models with various sources of datasets, state-of-th…

cs.CV202217 cited

Text-Adaptive Multiple Visual Prototype Matching for Video-Text Retrieval

Chengzhi Lin, Ancong Wu, Junwei Liang +4

Cross-modal retrieval between videos and texts has gained increasing research interest due to the rapid emergence of videos on the web. Generally, a video contains rich instance an…

cs.CV202242 cited

PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

Yuting Gao, Jinfeng Liu, Zihan Xu +4

Large-scale vision-language pre-training has achieved promising results on downstream tasks. Existing methods highly rely on the assumption that the image-text pairs crawled from t…