699 citations
- Centre National de la Recherche ScientifiqueFR24 papers
- Max Planck Institute for InformaticsDE23 papers
- German Research Centre for Artificial IntelligenceDE13 papers
- Karlsruhe Institute of TechnologyDE11 papers
- Universitat Autònoma de BarcelonaES11 papers
- Helmholtz Center for Information SecurityDE10 papers
- Leipzig UniversityDE10 papers
- Commissariat à l'Énergie Atomique et aux Énergies AlternativesFR9 papers
- Institute for Solid State Physics and OpticsHU9 papers
- Max Planck SocietyDE9 papers
- RWTH Aachen UniversityDE8 papers
- Universität HamburgDE8 papers
20 papers · 1 filter
Investigating Lexical Sharing in Multilingual Machine Translation for Indian Languages
Sonal Sannigrahi, Rachel Bawden
Multilingual language models have shown impressive cross-lingual transfer ability across a diverse set of languages and tasks. To improve the cross-lingual ability of these models,…
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
Isomorphic Cross-lingual Embeddings for Low-Resource Languages
Sonal Sannigrahi, Jesse Read
Cross-Lingual Word Embeddings (CLWEs) are a key component to transfer linguistic information learnt from higher-resource settings into lower-resource ones. Recent research in cross…
CLIN-X: pre-trained language models and a study on cross-task transfer for concept extraction in the clinical domain
Lukas Lange, Heike Adel, Jannik Strötgen +1
The field of natural language processing (NLP) has recently seen a large change towards using pre-trained language models for solving almost any task. Despite showing great improve…
Preventing Author Profiling through Zero-Shot Multilingual Back-Translation
David Ifeoluwa Adelani, Miaoran Zhang, Xiaoyu Shen +3
Documents as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g. their gender or ethnicity. Style transfer is an effective…
Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages
Dana Ruiter, Dietrich Klakow, Josef van Genabith +1
For most language combinations, parallel data is either scarce or simply unavailable. To address this, unsupervised machine translation (UMT) exploits large amounts of monolingual…