343 citations · 1.2k across the 55 of their papers we have counts for
24 papers · 1 filter
Attentive Multi-Layer Perceptron for Non-autoregressive Generation
Shuyang Jiang, Jun Zhang, Jiangtao Feng +2
Autoregressive~(AR) generation almost dominates sequence generation for its efficacy. Recently, non-autoregressive~(NAR) generation gains increasing popularity for its efficiency a…
Extrapolating Large Language Models to Non-English by Aligning Languages
Wenhao Zhu, Yunzhe Lv, Qingxiu Dong +6
Existing large language models show disparate capability across different languages, due to the imbalance in the training data. Their performances on English tasks are often strong…
Linearized Relative Positional Encoding
Zhen Qin, Weixuan Sun, Kaiyue Lu +6
Relative positional encoding is widely used in vanilla and linear transformers to represent positional information. However, existing encoding methods of a vanilla transformer are…
L-Eval: Instituting Standardized Evaluation for Long Context Language Models
Chenxin An, Shansan Gong, Ming Zhong +5
Recently, there has been growing interest in extending the context length of large language models (LLMs), aiming to effectively process long inputs of one turn or conversations wi…
Language Versatilists vs. Specialists: An Empirical Revisiting on Multilingual Transfer Ability
Jiacheng Ye, Xijia Tao, Lingpeng Kong
Multilingual transfer ability, which reflects how well the models fine-tuned on one source language can be applied to other languages, has been well studied in multilingual pre-tra…
INK: Injecting kNN Knowledge in Nearest Neighbor Machine Translation
Wenhao Zhu, Jingjing Xu, Shujian Huang +2
Neural machine translation has achieved promising results on many translation tasks. However, previous studies have shown that neural models induce a non-smooth representation spac…