18 citations · 20 across the 3 of their papers we have counts for
7 papers · 1 filter
MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation
Kaixin Cai, Pengzhen Ren, Yi Zhu +5
Recently, semantic segmentation models trained with image-level text supervision have shown promising results in challenging open-world scenarios. However, these models still face…
CapDet: Unifying Dense Captioning and Open-World Detection Pretraining
Yanxin Long, Youpeng Wen, Jianhua Han +5
Benefiting from large-scale vision-language pre-training on image-text pairs, open-world detection methods have shown superior generalization ability under the zero-shot or few-sho…
ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency
Pengzhen Ren, Changlin Li, Hang Xu +5
Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existi…
Beyond Fixation: Dynamic Window Visual Transformer
Pengzhen Ren, Changlin Li, Guangrun Wang +4
Recently, a surge of interest in visual transformers is to reduce the computational cost by limiting the calculation of self-attention to a local window. Most current work uses a f…
Unsupervised Person Re-Identification: A Systematic Survey of Challenges and Solutions
Xiangtan Lin, Pengzhen Ren, Chung-Hsing Yeh +3
Person re-identification (Re-ID) has been a significant research topic in the past decade due to its real-world applications and research significance. While supervised person Re-I…
Person Search Challenges and Solutions: A Survey
Xiangtan Lin, Pengzhen Ren, Yun Xiao +2
Person search has drawn increasing attention due to its real-world applications and research significance. Person search aims to find a probe person in a gallery of scene images wi…