55 citations · 56 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 55 cited
Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
An Yang, Junshu Pan, Junyang Lin +4
The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct…
cs.CV2021★ 1 cited
Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization
Litian Zhang, Xiaoming Zhang, Junshu Pan +1
Multimodal summarization with multimodal output (MSMO) generates a summary with both textual and visual content. Multimodal news report contains heterogeneous contents, which makes…