25 citations · 25 across the 1 of their papers we have counts for
3 papers
cs.CL2021
M6: A Chinese Multimodal Pretrainer
Junyang Lin, Rui Men, An Yang +22
In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We pro…
cs.CV2020★ 25 cited
Poet: Product-oriented Video Captioner for E-commerce
Shengyu Zhang, Ziqi Tan, Jin Yu +6
In e-commerce, a growing number of user-generated videos are used for product promotion. How to generate video descriptions that narrate the user-preferred product characteristics…
cs.CL2020
InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
Junyang Lin, An Yang, Yichang Zhang +3
Multi-modal pretraining for learning high-level multi-modal representation is a further step towards deep learning and artificial intelligence. In this work, we propose a novel mod…