activity
20182022
most citedDAPPLE: A Pipelined Data Parallel Approach for Training Large Models

29 citations · 83 across the 13 of their papers we have counts for

collaborators
Showing 2021Show all

5 papers · 1 filter

cs.LG202118 cited

M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Junyang Lin, An Yang, Jinze Bai +9

Recent expeditious developments in deep learning algorithms, distributed training, and even hardware design for large models have enabled training extreme-scale models, say GPT-3 a…

cs.AI2021

Explicit Semantic Cross Feature Learning via Pre-trained Graph Neural Networks for CTR Prediction

Feng Li, Bencheng Yan, Qingqing Long +4

Cross features play an important role in click-through rate (CTR) prediction. Most of the existing methods adopt a DNN-based model to capture the cross features in an implicit mann…

cs.IR2021

Towards a Better Tradeoff between Effectiveness and Efficiency in Pre-Ranking: A Learnable Feature Selection based Approach

Xu Ma, Pengjie Wang, Hui Zhao +6

In real-world search, recommendation, and advertising systems, the multi-stage ranking architecture is commonly adopted. Such architecture usually consists of matching, pre-ranking…

cs.LG2021

M6-T: Exploring Sparse Expert Models and Beyond

An Yang, Junyang Lin, Rui Men +12

Mixture-of-Experts (MoE) models can achieve promising results with outrageous large amount of parameters but constant computation cost, and thus it has become a trend in model scal…

cs.CL2021

M6: A Chinese Multimodal Pretrainer

Junyang Lin, Rui Men, An Yang +22

In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We pro…