2 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 2 cited
FashionSAP: Symbols and Attributes Prompt for Fine-grained Fashion Vision-Language Pre-training
Yunpeng Han, Lisai Zhang, Qingcai Chen +4
Fashion vision-language pre-training models have shown efficacy for a wide range of downstream tasks. However, general vision-language pre-training models pay less attention to fin…
cs.CV2023★ 1 cited
Refined Vision-Language Modeling for Fine-grained Multi-modal Pre-training
Lisai Zhang, Qingcai Chen, Zhijian Chen +3
Fine-grained supervision based on object annotations has been widely used for vision and language pre-training (VLP). However, in real-world application scenarios, aligned multi-mo…
cs.CV2021★ 1 cited
VLDeformer: Vision-Language Decomposed Transformer for Fast Cross-Modal Retrieval
Lisai Zhang, Hongfa Wu, Qingcai Chen +6
Cross-model retrieval has emerged as one of the most important upgrades for text-only search engines (SE). Recently, with powerful representation for pairwise text-image inputs via…