57 citations · 57 across the 3 of their papers we have counts for
2 papers
cs.LG2022★ 57 cited
Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism
Xupeng Miao, Yujie Wang, Youhe Jiang +4
Transformer models have achieved state-of-the-art performance on various domains of applications and gradually becomes the foundations of the advanced large deep learning (DL) mode…
cs.LG2021★ 57 cited
HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework
Xupeng Miao, Hailin Zhang, Yining Shi +4
Embedding models have been an effective learning paradigm for high-dimensional data. However, one open issue of embedding models is that their representations (latent factors) ofte…