output
20022024
most citedNWChem: Past, Present, and Future

699 citations

Showing cs.CLShow all

20 papers · 1 filter

cs.CL20231 cited

Investigating Lexical Sharing in Multilingual Machine Translation for Indian Languages

Sonal Sannigrahi, Rachel Bawden

Multilingual language models have shown impressive cross-lingual transfer ability across a diverse set of languages and tasks. To improve the cross-lingual ability of these models,…

cs.CL202365 cited

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Hugo Laurençon, Lucile Saulnier, Thomas Wang +51

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…

cs.CL20222 cited

Isomorphic Cross-lingual Embeddings for Low-Resource Languages

Sonal Sannigrahi, Jesse Read

Cross-Lingual Word Embeddings (CLWEs) are a key component to transfer linguistic information learnt from higher-resource settings into lower-resource ones. Recent research in cross…

cs.CL202117 cited

CLIN-X: pre-trained language models and a study on cross-task transfer for concept extraction in the clinical domain

Lukas Lange, Heike Adel, Jannik Strötgen +1

The field of natural language processing (NLP) has recently seen a large change towards using pre-trained language models for solving almost any task. Despite showing great improve…

cs.CL2021

Preventing Author Profiling through Zero-Shot Multilingual Back-Translation

David Ifeoluwa Adelani, Miaoran Zhang, Xiaoyu Shen +3

Documents as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g. their gender or ethnicity. Style transfer is an effective…

cs.CL20212 cited

Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages

Dana Ruiter, Dietrich Klakow, Josef van Genabith +1

For most language combinations, parallel data is either scarce or simply unavailable. To address this, unsupervised machine translation (UMT) exploits large amounts of monolingual…