8 citations · 17 across the 6 of their papers we have counts for
6 papers · 1 filter
Extrapolating Large Language Models to Non-English by Aligning Languages
Wenhao Zhu, Yunzhe Lv, Qingxiu Dong +6
Existing large language models show disparate capability across different languages, due to the imbalance in the training data. Their performances on English tasks are often strong…
TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer
Zhen Qin, Dong Li, Weigao Sun +8
We present TransNormerLLM, the first linear attention-based Large Language Model (LLM) that outperforms conventional softmax attention-based models in terms of both accuracy and ef…
Extrapolating Multilingual Understanding Models as Multilingual Generators
Bohong Wu, Fei Yuan, Hai Zhao +2
Multilingual understanding models (or encoder-based), pre-trained via masked language modeling, have achieved promising results on many language understanding tasks (e.g., mBERT).…
Simpson's Bias in NLP Training
Fei Yuan, Longtu Zhang, Huang Bojun +1
In most machine learning tasks, we evaluate a model on a given data population by measuring a population-level metric . Examples of such evaluation metric inclu…
Reinforced Multi-Teacher Selection for Knowledge Distillation
Fei Yuan, Linjun Shou, Jian Pei +4
In natural language processing (NLP) tasks, slow inference speed and huge footprints in GPU usage remain the bottleneck of applying pre-trained deep models in production. As a popu…
Enhancing Answer Boundary Detection for Multilingual Machine Reading Comprehension
Fei Yuan, Linjun Shou, Xuanyu Bai +5
Multilingual pre-trained models could leverage the training data from a rich source language (such as English) to improve performance on low resource languages. However, the transf…