2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2023
Accelerating Large Batch Training via Gradient Signal to Noise Ratio (GSNR)
Guo-qing Jiang, Jinlong Liu, Zixiang Ding +2
As models for nature language processing (NLP), computer vision (CV) and recommendation systems (RS) require surging computation, a large number of GPUs/TPUs are paralleled as a la…
cs.CL2022★ 2 cited
SKDBERT: Compressing BERT via Stochastic Knowledge Distillation
Zixiang Ding, Guoqing Jiang, Shuai Zhang +2
In this paper, we propose Stochastic Knowledge Distillation (SKD) to obtain compact BERT-style language model dubbed SKDBERT. In each iteration, SKD samples a teacher model from a…