36 citations · 88 across the 6 of their papers we have counts for
3 papers · 1 filter
Improving Text Generation with Student-Forcing Optimal Transport
Guoyin Wang, Chunyuan Li, Jianqiao Li +10
Neural language models are often trained with maximum likelihood estimation (MLE), where the next word is generated conditioned on the ground-truth word tokens. During testing, how…
Ouroboros: On Accelerating Training of Transformer-Based Language Models
Qian Yang, Zhouyuan Huo, Wenlin Wang +2
Language models are essential for natural language processing (NLP) tasks, such as machine translation and text summarization. Remarkable performance has been demonstrated recently…
Learning Compressed Sentence Representations for On-Device Text Processing
Dinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman +5
Vector representations of sentences, trained on massive text corpora, are widely used as generic sentence embeddings across a variety of NLP problems. The learned representations a…