activity
20182022
most citedMinimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation

41 citations · 146 across the 20 of their papers we have counts for

collaborators

31 papers

cs.CL20211 cited

Guiding Teacher Forcing with Seer Forcing for Neural Machine Translation

Yang Feng, Shuhao Gu, Dengji Guo +2

Although teacher forcing has become the main training paradigm for neural machine translation, it usually makes predictions only conditioned on past information, and hence lacks gl…

cs.CL2021

GTM: A Generative Triple-Wise Model for Conversational Question Generation

Lei Shen, Fandong Meng, Jinchao Zhang +2

Generating some appealing questions in open-domain conversations is an effective way to improve human-machine interactions and lead the topic to a broader or deeper direction. To a…

cs.CL20211 cited

Addressing Inquiries about History: An Efficient and Practical Framework for Evaluating Open-domain Chatbot Consistency

Zekang Li, Jinchao Zhang, Zhengcong Fei +2

A good open-domain chatbot should avoid presenting contradictory responses about facts or opinions in a conversational session, known as its consistency capacity. However, evaluati…

cs.CL20215 cited

Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue Utterances

Zekang Li, Jinchao Zhang, Zhengcong Fei +2

Nowadays, open-domain dialogue models can generate acceptable responses according to the historical context based on the large-scale pre-trained language models. However, they gene…

cs.CL20212 cited

Sequence-Level Training for Non-Autoregressive Neural Machine Translation

Chenze Shao, Yang Feng, Jinchao Zhang +2

In recent years, Neural Machine Translation (NMT) has achieved notable results in various translation tasks. However, the word-by-word generation manner determined by the autoregre…

cs.CL2021

Modeling Coverage for Non-Autoregressive Neural Machine Translation

Yong Shan, Yang Feng, Chenze Shao

Non-Autoregressive Neural Machine Translation (NAT) has achieved significant inference speedup by generating all tokens simultaneously. Despite its high efficiency, NAT usually suf…