90 citations · 150 across the 9 of their papers we have counts for
13 papers · 1 filter
GRIN: GRadient-INformed MoE
Liyuan Liu, Young Jin Kim, Shuohang Wang +14
Mixture-of-Experts (MoE) models scale more effectively than dense models due to sparse computation through expert routing, selectively activating only a small subset of expert modu…
Open-domain Question Answering via Chain of Reasoning over Heterogeneous Knowledge
Kaixin Ma, Hao Cheng, Xiaodong Liu +2
We propose a novel open-domain question answering (ODQA) framework for answering single/multi-hop questions across heterogeneous knowledge sources. The key novelty of our method is…
A Survey of Knowledge-Intensive NLP with Pre-Trained Language Models
Da Yin, Li Dong, Hao Cheng +4
With the increasing of model capacity brought by pre-trained language models, there emerges boosting needs for more knowledgeable natural language processing (NLP) models with adva…
CLUES: Few-Shot Learning Evaluation in Natural Language Understanding
Subhabrata Mukherjee, Xiaodong Liu, Guoqing Zheng +6
Most recent progress in natural language understanding (NLU) has been driven, in part, by benchmarks such as GLUE, SuperGLUE, SQuAD, etc. In fact, many NLU models have now matched…
Dialogue State Tracking with a Language Model using Schema-Driven Prompting
Chia-Hsuan Lee, Hao Cheng, Mari Ostendorf
Task-oriented conversational systems often use dialogue state tracking to represent the user's intentions, which involves filling in values of pre-defined slots. Many approaches ha…
Targeted Adversarial Training for Natural Language Understanding
Lis Pereira, Xiaodong Liu, Hao Cheng +3
We present a simple yet effective Targeted Adversarial Training (TAT) algorithm to improve adversarial training for natural language understanding. The key idea is to introspect cu…