activity
20122022
most citedCyclical Annealing Schedule: A Simple Approach to Mitigating KL Vanishing

169 citations · 1.5k across the 75 of their papers we have counts for

collaborators
Showing cs.CLShow all

25 papers · 1 filter

cs.CL2022

Pseudo-OOD training for robust language models

Dhanasekar Sundararaman, Nikhil Mehta, Lawrence Carin

While pre-trained large-scale deep models have garnered attention as an important topic for many downstream natural language processing (NLP) tasks, such models often make unreliab…

cs.CL202151 cited

FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders

Pengyu Cheng, Weituo Hao, Siyang Yuan +2

Pretrained text encoders, such as BERT, have been applied increasingly in various natural language processing (NLP) tasks, and have recently demonstrated significant performance ga…

cs.CL2021159 cited

What Makes Good In-Context Examples for GPT-?

Jiachang Liu, Dinghan Shen, Yizhe Zhang +3

GPT- has attracted lots of attention due to its superior performance across a wide range of NLP tasks, especially with its powerful and versatile in-context few-shot learning ab…

cs.CL202030 cited

MixKD: Towards Efficient Distillation of Large-scale Language Models

Kevin J Liang, Weituo Hao, Dinghan Shen +4

Large-scale language models have recently demonstrated impressive empirical performance. Nevertheless, the improved results are attained at the price of bigger models, more power c…

cs.CL2020

Improving Text Generation with Student-Forcing Optimal Transport

Guoyin Wang, Chunyuan Li, Jianqiao Li +10

Neural language models are often trained with maximum likelihood estimation (MLE), where the next word is generated conditioned on the ground-truth word tokens. During testing, how…

cs.CL202022 cited

Graph Optimal Transport for Cross-Domain Alignment

Liqun Chen, Zhe Gan, Yu Cheng +3

Cross-domain alignment between two sets of entities (e.g., objects in an image, words in a sentence) is fundamental to both computer vision and natural language processing. Existin…