activity
20142024
most citedARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design

27 citations · 205 across the 59 of their papers we have counts for

collaborators
Showing cs.CVShow all

17 papers · 1 filter

cs.CV20232 cited

UniDiff: Advancing Vision-Language Models with Generative and Discriminative Learning

Xiao Dong, Runhui Huang, Xiaoyong Wei +4

Recent advances in vision-language pre-training have enabled machines to perform better in multimodal object discrimination (e.g., image-text semantic alignment) and image synthesi…

cs.CV20232 cited

DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment

Lewei Yao, Jianhua Han, Xiaodan Liang +4

This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike…

cs.CV20233 cited

CLIP: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data

Yihan Zeng, Chenhan Jiang, Jiageng Mao +7

Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. Howeve…

cs.CV20239 cited

GP-VTON: Towards General Purpose Virtual Try-on via Collaborative Local-Flow Global-Parsing Learning

Zhenyu Xie, Zaiyu Huang, Xin Dong +5

Image-based Virtual Try-ON aims to transfer an in-shop garment onto a specific person. Existing methods employ a global warping module to model the anisotropic deformation for diff…

cs.CV20235 cited

Dynamic Graph Enhanced Contrastive Learning for Chest X-ray Report Generation

Mingjie Li, Bingqian Lin, Zicong Chen +3

Automatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced da…

cs.CV20232 cited

CapDet: Unifying Dense Captioning and Open-World Detection Pretraining

Yanxin Long, Youpeng Wen, Jianhua Han +5

Benefiting from large-scale vision-language pre-training on image-text pairs, open-world detection methods have shown superior generalization ability under the zero-shot or few-sho…