2 papers
cs.CL2021
M6: A Chinese Multimodal Pretrainer
Junyang Lin, Rui Men, An Yang +22
In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We pro…
cs.SI2020
A Multi-Semantic Metapath Model for Large Scale Heterogeneous Network Representation Learning
Xuandong Zhao, Jinbao Xue, Jin Yu +2
Network Embedding has been widely studied to model and manage data in a variety of real-world applications. However, most existing works focus on networks with single-typed nodes o…