57 citations · 72 across the 3 of their papers we have counts for
3 papers
cs.LG2022★ 57 cited
Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism
Xupeng Miao, Yujie Wang, Youhe Jiang +4
Transformer models have achieved state-of-the-art performance on various domains of applications and gradually becomes the foundations of the advanced large deep learning (DL) mode…
cs.CV2022★ 14 cited
Knowledge Distillation with the Reused Teacher Classifier
Defang Chen, Jian-Ping Mei, Hailin Zhang +3
Knowledge distillation aims to compress a powerful yet cumbersome teacher model into a lightweight student model without much sacrifice of performance. For this purpose, various ap…
cs.LG2022★ 1 cited
Confidence-Aware Multi-Teacher Knowledge Distillation
Hailin Zhang, Defang Chen, Can Wang
Knowledge distillation is initially introduced to utilize additional supervision from a single teacher model for the student model training. To boost the student performance, some…