12 citations · 12 across the 1 of their papers we have counts for
3 papers
cs.CL2020
Taking Notes on the Fly Helps BERT Pre-training
Qiyu Wu, Chen Xing, Yatao Li +3
How to make unsupervised language pre-training more efficient and less resource-intensive is an important research direction in NLP. In this paper, we focus on improving the effici…
cs.LG2020
On Layer Normalization in the Transformer Architecture
Ruibin Xiong, Yunchang Yang, Di He +7
The Transformer is widely used in natural language processing tasks. To train a Transformer however, one usually needs a carefully designed learning rate warm-up stage, which is sh…
cs.LG2019★ 12 cited
Distance-Based Learning from Errors for Confidence Calibration
Chen Xing, Sercan Arik, Zizhao Zhang +1
Deep neural networks (DNNs) are poorly calibrated when trained in conventional ways. To improve confidence calibration of DNNs, we propose a novel training method, distance-based l…