27 citations · 205 across the 59 of their papers we have counts for
17 papers · 1 filter
UniDiff: Advancing Vision-Language Models with Generative and Discriminative Learning
Xiao Dong, Runhui Huang, Xiaoyong Wei +4
Recent advances in vision-language pre-training have enabled machines to perform better in multimodal object discrimination (e.g., image-text semantic alignment) and image synthesi…
DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment
Lewei Yao, Jianhua Han, Xiaodan Liang +4
This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike…
CLIP: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data
Yihan Zeng, Chenhan Jiang, Jiageng Mao +7
Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. Howeve…
GP-VTON: Towards General Purpose Virtual Try-on via Collaborative Local-Flow Global-Parsing Learning
Zhenyu Xie, Zaiyu Huang, Xin Dong +5
Image-based Virtual Try-ON aims to transfer an in-shop garment onto a specific person. Existing methods employ a global warping module to model the anisotropic deformation for diff…
Dynamic Graph Enhanced Contrastive Learning for Chest X-ray Report Generation
Mingjie Li, Bingqian Lin, Zicong Chen +3
Automatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced da…
CapDet: Unifying Dense Captioning and Open-World Detection Pretraining
Yanxin Long, Youpeng Wen, Jianhua Han +5
Benefiting from large-scale vision-language pre-training on image-text pairs, open-world detection methods have shown superior generalization ability under the zero-shot or few-sho…