97 citations · 111 across the 4 of their papers we have counts for
5 papers
ODE Transformer: An Ordinary Differential Equation-Inspired Model for Neural Machine Translation
Bei Li, Quan Du, Tao Zhou +4
It has been found that residual networks are an Euler discretization of solutions to Ordinary Differential Equations (ODEs). In this paper, we explore a deeper relationship between…
Learning Light-Weight Translation Models from Deep Transformer
Bei Li, Ziyang Wang, Hui Liu +4
Recently, deep models have shown tremendous improvements in neural machine translation (NMT). However, systems of this kind are computationally expensive and memory intensive. In t…
Shallow-to-Deep Training for Neural Machine Translation
Bei Li, Ziyang Wang, Hui Liu +5
Deep encoders have been proven to be effective in improving neural machine translation (NMT) systems, but training an extremely deep encoder is time consuming. Moreover, why deep m…
Does Multi-Encoder Help? A Case Study on Context-Aware Neural Machine Translation
Bei Li, Hui Liu, Ziyang Wang +5
In encoder-decoder neural models, multiple encoders are in general used to represent the contextual information in addition to the individual sentence. In this paper, we investigat…
Learning Deep Transformer Models for Machine Translation
Qiang Wang, Bei Li, Tong Xiao +4
Transformer is the state-of-the-art model in recent machine translation evaluations. Two strands of research are promising to improve models of this kind: the first uses wide netwo…