activity
20182022
most citedYou Only Compress Once: Towards Effective and Elastic BERT Compression via Exploit-Explore Stochastic Nature Gradient

11 citations · 38 across the 9 of their papers we have counts for

collaborators
Showing 2020Show all

5 papers · 1 filter

cs.CL2020

EasyTransfer -- A Simple and Scalable Deep Transfer Learning Platform for NLP Applications

Minghui Qiu, Peng Li, Chengyu Wang +8

The literature has witnessed the success of leveraging Pre-trained Language Models (PLMs) and Transfer Learning (TL) algorithms to a wide range of Natural Language Processing (NLP)…

cs.SD20208 cited

INT8 Winograd Acceleration for Conv1D Equipped ASR Models Deployed on Mobile Devices

Yiwu Yao, Yuchao Li, Chengyu Wang +8

The intensive computation of Automatic Speech Recognition (ASR) models obstructs them from being deployed on mobile devices. In this paper, we present a novel quantized Winograd op…

cs.DC20205 cited

Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads

Siyu Wang, Yi Rong, Shiqing Fan +6

The last decade has witnessed growth in the computational requirements for training deep neural networks. Current approaches (e.g., data/model parallelism, pipeline parallelism) pa…

cs.DC20201 cited

DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging

Qinggang Zhou, Yawen Zhang, Pengcheng Li +4

The state-of-the-art deep learning algorithms rely on distributed training systems to tackle the increasing sizes of models and training data sets. Minibatch stochastic gradient de…

cs.LG2020

Auto-Ensemble: An Adaptive Learning Rate Scheduling based Deep Learning Model Ensembling

Jun Yang, Fei Wang

Ensembling deep learning models is a shortcut to promote its implementation in new scenarios, which can avoid tuning neural networks, losses and training algorithms from scratch. H…