activity
20162025
most citedDomain-specific Continued Pretraining of Language Models for Capturing Long Context in Mental Health

16 citations · 30 across the 12 of their papers we have counts for

collaborators

12 papers

cs.CL2025

DocHPLT: A Massively Multilingual Document-Level Translation Dataset

Dayyán O'Brien, Bhavitvya Malik, Ona de Gibert +3

Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of docume…

cs.CL2025

SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes

Raúl Vázquez, Timothee Mickus, Elaine Zosa +15

We present the Mu-SHROOM shared task which is focused on detecting hallucinations and other overgeneration mistakes in the output of instruction-tuned large language models (LLMs).…

cs.CL2024

A Comparison of Language Modeling and Translation as Multilingual Pretraining Objectives

Zihao Li, Shaoxiong Ji, Timothee Mickus +2

Pretrained language models (PLMs) display impressive performances and have captured the attention of the NLP community. Establishing best practices in pretraining has, therefore, b…

cs.CL2024

SemEval-2024 Shared Task 6: SHROOM, a Shared-task on Hallucinations and Related Observable Overgeneration Mistakes

Timothee Mickus, Elaine Zosa, Raúl Vázquez +5

This paper presents the results of the SHROOM, a shared task focused on detecting hallucinations: outputs from natural language generation (NLG) systems that are fluent, yet inaccu…

cs.CL2024

Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?

Shaoxiong Ji, Timothee Mickus, Vincent Segonne +1

Multilingual pretraining and fine-tuning have remarkably succeeded in various natural language processing tasks. Transferring representations from one language to another is especi…

cs.CL20246 cited

A New Massive Multilingual Dataset for High-Performance Language Technologies

Ona de Gibert, Graeme Nail, Nikolay Arefyev +10

We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from…