most citedSharing Attention Weights for Fast Transformer

7 citations · 13 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2020

Towards Fully 8-bit Integer Inference for the Transformer Model

Ye Lin, Yanyang Li, Tengbo Liu +3

8-bit integer inference, as a promising direction in reducing both the latency and storage of deep neural networks, has made great progress recently. On the other hand, previous sy…

cs.LG20201 cited

Learning Architectures from an Extended Search Space for Language Modeling

Yinqiao Li, Chi Hu, Yuhao Zhang +6

Neural architecture search (NAS) has advanced significantly in recent years but most NAS systems restrict search to learning architectures of a recurrent or convolutional cell. In…

cs.CL20201 cited

Does Multi-Encoder Help? A Case Study on Context-Aware Neural Machine Translation

Bei Li, Hui Liu, Ziyang Wang +5

In encoder-decoder neural models, multiple encoders are in general used to represent the contextual information in addition to the individual sentence. In this paper, we investigat…

cs.CL20204 cited

Neural Machine Translation with Joint Representation

Yanyang Li, Qiang Wang, Tong Xiao +2

Though early successes of Statistical Machine Translation (SMT) systems are attributed in part to the explicit modelling of the interaction between any two source and target units,…

cs.CL20197 cited

Sharing Attention Weights for Fast Transformer

Tong Xiao, Yinqiao Li, Jingbo Zhu +2

Recently, the Transformer machine translation system has shown strong results by stacking attention layers on both the source and target-language sides. But the inference of this m…