6 papers
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
Peiqin Lin, Chenyang Lyu, Wenjiang Luo +22
Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prio…
EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models
Shaoxiong Ji, Zihao Li, Jaakko Paavola +7
In this work, we introduce EMMA-500, a large-scale multilingual language model continue-trained on texts across 546 languages designed for enhanced multilingual performance, focusi…
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models
Hengyu Luo, Zihao Li, Joseph Attieh +12
Large language models (LLMs) are advancing at an unprecedented pace globally, with regions increasingly adopting these models for applications in their primary language. Evaluation…
Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on Manchu
Renhao Pei, Yihong Liu, Peiqin Lin +2
In-context machine translation (MT) with large language models (LLMs) is a promising approach for low-resource MT, as it can readily take advantage of linguistic resources such as…
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models
Peiqin Lin, André F. T. Martins, Hinrich Schütze
Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models, improving performance in both bilingual tasks, e.g., mac…
XAMPLER: Learning to Retrieve Cross-Lingual In-Context Examples
Peiqin Lin, André F. T. Martins, Hinrich Schütze
Recent studies indicate that leveraging off-the-shelf or fine-tuned retrievers, capable of retrieving relevant in-context examples tailored to the input query, enhances few-shot in…