activity
20202023
most citedExtrapolating Large Language Models to Non-English by Aligning Languages

8 citations · 17 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL20238 cited

Extrapolating Large Language Models to Non-English by Aligning Languages

Wenhao Zhu, Yunzhe Lv, Qingxiu Dong +6

Existing large language models show disparate capability across different languages, due to the imbalance in the training data. Their performances on English tasks are often strong…

cs.CL2023

TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer

Zhen Qin, Dong Li, Weigao Sun +8

We present TransNormerLLM, the first linear attention-based Large Language Model (LLM) that outperforms conventional softmax attention-based models in terms of both accuracy and ef…

cs.CL20231 cited

Extrapolating Multilingual Understanding Models as Multilingual Generators

Bohong Wu, Fei Yuan, Hai Zhao +2

Multilingual understanding models (or encoder-based), pre-trained via masked language modeling, have achieved promising results on many language understanding tasks (e.g., mBERT).…

cs.CL2021

Simpson's Bias in NLP Training

Fei Yuan, Longtu Zhang, Huang Bojun +1

In most machine learning tasks, we evaluate a model on a given data population by measuring a population-level metric . Examples of such evaluation metric inclu…

cs.CL20202 cited

Reinforced Multi-Teacher Selection for Knowledge Distillation

Fei Yuan, Linjun Shou, Jian Pei +4

In natural language processing (NLP) tasks, slow inference speed and huge footprints in GPU usage remain the bottleneck of applying pre-trained deep models in production. As a popu…

cs.CL20206 cited

Enhancing Answer Boundary Detection for Multilingual Machine Reading Comprehension

Fei Yuan, Linjun Shou, Xuanyu Bai +5

Multilingual pre-trained models could leverage the training data from a rich source language (such as English) to improve performance on low resource languages. However, the transf…