11 citations · 43 across the 12 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2022★ 3 cited
Mixture of Attention Heads: Selecting Attention Heads Per Token
Xiaofeng Zhang, Yikang Shen, Zeyu Huang +3
Mixture-of-Experts (MoE) networks have been proposed as an efficient way to scale up model capacity and implement conditional computing. However, the study of MoE components mostly…
cs.CL2022★ 1 cited
Entity Aware Syntax Tree Based Data Augmentation for Natural Language Understanding
Jiaxing Xu, Jianbin Cui, Jiangneng Li +2
Understanding the intention of the users and recognizing the semantic entities from their sentences, aka natural language understanding (NLU), is the upstream task of many natural…
cs.CL2018
Why Do Neural Response Generation Models Prefer Universal Replies?
Bowen Wu, Nan Jiang, Zhifeng Gao +6
Recent advances in sequence-to-sequence learning reveal a purely data-driven approach to the response generation task. Despite its diverse applications, existing neural models are…