1.2k citations · 1.3k across the 7 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2022★ 1.2k cited
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre +32
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we expl…
cs.LG2021★ 3 cited
Speeding up Deep Model Training by Sharing Weights and Then Unsharing
Shuo Yang, Le Hou, Xiaodan Song +2
We propose a simple and efficient approach for training the BERT model. Our approach exploits the special structure of BERT that contains a stack of repeated modules (i.e., transfo…