86 citations · 106 across the 3 of their papers we have counts for
10 papers · 1 filter
A Comparison of Language Modeling and Translation as Multilingual Pretraining Objectives
Zihao Li, Shaoxiong Ji, Timothee Mickus +2
Pretrained language models (PLMs) display impressive performances and have captured the attention of the NLP community. Establishing best practices in pretraining has, therefore, b…
Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?
Shaoxiong Ji, Timothee Mickus, Vincent Segonne +1
Multilingual pretraining and fine-tuning have remarkably succeeded in various natural language processing tasks. Transferring representations from one language to another is especi…
A New Massive Multilingual Dataset for High-Performance Language Technologies
Ona de Gibert, Graeme Nail, Nikolay Arefyev +10
We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from…
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki
Timothee Mickus, Stig-Arne Grönroos, Joseph Attieh +7
NLP in the age of monolithic large language models is approaching its limits in terms of size and information that can be handled. The trend goes to modularization, a necessary ste…
MaLA-500: Massive Language Adaptation of Large Language Models
Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann +2
Large language models (LLMs) have advanced the state of the art in natural language processing. However, their predominant design for English or a limited set of languages creates…
Content Reduction, Surprisal and Information Density Estimation for Long Documents
Shaoxiong Ji, Wei Sun, Pekka Marttinen
Many computational linguistic methods have been proposed to study the information content of languages. We consider two interesting research questions: 1) how is information distri…