14 citations · 18 across the 5 of their papers we have counts for
Showing 2021 · cs.CLShow all
3 papers · 2 filters
cs.CL2021★ 1 cited
FastSeq: Make Sequence Generation Faster
Yu Yan, Fei Hu, Jiusheng Chen +7
Transformer-based models have made tremendous impacts in natural language generation. However the inference speed is a bottleneck due to large model size and intensive computing in…
cs.CL2021★ 2 cited
EL-Attention: Memory Efficient Lossless Attention for Generation
Yu Yan, Jiusheng Chen, Weizhen Qi +4
Transformer model with multi-head attention requires caching intermediate results for efficient inference in generation tasks. However, cache brings new memory-related costs and pr…
cs.CL2021
ProphetNet-X: Large-Scale Pre-training Models for English, Chinese, Multi-lingual, Dialog, and Code Generation
Weizhen Qi, Yeyun Gong, Yu Yan +9
Now, the pre-training technique is ubiquitous in natural language processing field. ProphetNet is a pre-training based natural language generation method which shows powerful perfo…