14 citations · 20 across the 7 of their papers we have counts for
7 papers
UniDiff: Advancing Vision-Language Models with Generative and Discriminative Learning
Xiao Dong, Runhui Huang, Xiaoyong Wei +4
Recent advances in vision-language pre-training have enabled machines to perform better in multimodal object discrimination (e.g., image-text semantic alignment) and image synthesi…
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
Xiao Wang, Guangyao Chen, Guangwu Qian +5
With the urgent demand for generalized deep models, many pre-trained big models are proposed, such as BERT, ViT, GPT, etc. Inspired by the success of these models in single domains…
Entity-Graph Enhanced Cross-Modal Pretraining for Instance-level Product Retrieval
Xiao Dong, Xunlin Zhan, Yunchao Wei +5
Our goal in this research is to study a more realistic environment in which we can conduct weakly-supervised multi-modal instance-level product retrieval for fine-grained product c…
Indicative Image Retrieval: Turning Blackbox Learning into Grey
Xulu Zhang, Zhenqun Yang, Hao Tian +2
Deep learning became the game changer for image retrieval soon after it was introduced. It promotes the feature extraction (by representation learning) as the core of image retriev…
Deep learning-based person re-identification methods: A survey and outlook of recent works
Zhangqiang Ming, Min Zhu, Xiangkun Wang +5
In recent years, with the increasing demand for public safety and the rapid development of intelligent surveillance networks, person re-identification (Re-ID) has become one of the…
Global-Local Dynamic Feature Alignment Network for Person Re-Identification
Zhangqiang Ming, Yong Yang, Xiaoyong Wei +4
The misalignment of human images caused by bounding box detection errors or partial occlusions is one of the main challenges in person Re-Identification (Re-ID) tasks. Previous loc…