activity
20192022
most citedM6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

18 citations · 39 across the 4 of their papers we have counts for

collaborators

8 papers

cs.DC2022

PICASSO: Unleashing the Potential of GPU-centric Training for Wide-and-deep Recommender Systems

Yuanxing Zhang, Langshi Chen, Siran Yang +12

The development of personalized recommendation has significantly improved the accuracy of information matching and the revenue of e-commerce platforms. Recently, it has 2 trends: 1…

cs.LG202118 cited

M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Junyang Lin, An Yang, Jinze Bai +9

Recent expeditious developments in deep learning algorithms, distributed training, and even hardware design for large models have enabled training extreme-scale models, say GPT-3 a…

cs.LG2021

M6-T: Exploring Sparse Expert Models and Beyond

An Yang, Junyang Lin, Rui Men +12

Mixture-of-Experts (MoE) models can achieve promising results with outrageous large amount of parameters but constant computation cost, and thus it has become a trend in model scal…

cs.CL2021

M6: A Chinese Multimodal Pretrainer

Junyang Lin, Rui Men, An Yang +22

In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We pro…

cs.CV20206 cited

VEGA: Towards an End-to-End Configurable AutoML Pipeline

Bochao Wang, Hang Xu, Jiajin Zhang +21

Automated Machine Learning (AutoML) is an important industrial solution for automatic discovery and deployment of the machine learning models. However, designing an integrated Auto…

cs.CL2020

EasyTransfer -- A Simple and Scalable Deep Transfer Learning Platform for NLP Applications

Minghui Qiu, Peng Li, Chengyu Wang +8

The literature has witnessed the success of leveraging Pre-trained Language Models (PLMs) and Transfer Learning (TL) algorithms to a wide range of Natural Language Processing (NLP)…