86 citations · 120 across the 11 of their papers we have counts for
11 papers
Graph2text or Graph2token: A Perspective of Large Language Models for Graph Learning
Shuo Yu, Yingbo Wang, Ruolin Li +7
Graphs are data structures used to represent irregular networks and are prevalent in numerous real-world applications. Previous methods directly model graph structures and achieve…
A Comparison of Language Modeling and Translation as Multilingual Pretraining Objectives
Zihao Li, Shaoxiong Ji, Timothee Mickus +2
Pretrained language models (PLMs) display impressive performances and have captured the attention of the NLP community. Establishing best practices in pretraining has, therefore, b…
Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?
Shaoxiong Ji, Timothee Mickus, Vincent Segonne +1
Multilingual pretraining and fine-tuning have remarkably succeeded in various natural language processing tasks. Transferring representations from one language to another is especi…
A New Massive Multilingual Dataset for High-Performance Language Technologies
Ona de Gibert, Graeme Nail, Nikolay Arefyev +10
We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from…
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki
Timothee Mickus, Stig-Arne Grönroos, Joseph Attieh +7
NLP in the age of monolithic large language models is approaching its limits in terms of size and information that can be handled. The trend goes to modularization, a necessary ste…
MaLA-500: Massive Language Adaptation of Large Language Models
Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann +2
Large language models (LLMs) have advanced the state of the art in natural language processing. However, their predominant design for English or a limited set of languages creates…