75 citations · 89 across the 5 of their papers we have counts for
7 papers
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
Pedro Javier Ortiz Suárez, Laurent Romary, Benoît Sagot
We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) fo…
Establishing a New State-of-the-Art for French Named Entity Recognition
Pedro Javier Ortiz Suárez, Yoann Dupont, Benjamin Muller +2
The French TreeBank developed at the University Paris 7 is the main source of morphosyntactic and syntactic annotations for French. However, it does not include explicit informatio…
ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting Transformations
Fernando Alva-Manchego, Louis Martin, Antoine Bordes +3
In order to simplify a sentence, human editors perform multiple rewriting transformations: they split it into several shorter sentences, paraphrase words (i.e. replacing complex wo…
Can Multilingual Language Models Transfer to an Unseen Dialect? A Case Study on North African Arabizi
Benjamin Muller, Benoit Sagot, Djamé Seddah
Building natural language processing systems for non standardized and low resource languages is a difficult challenge. The recent success of large-scale multilingual pretrained lan…
CamemBERT: a Tasty French Language Model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez +5
Pretrained language models are now ubiquitous in Natural Language Processing. Despite their success, most available models have either been trained on English data or on the concat…
Controllable Sentence Simplification
Louis Martin, Benoît Sagot, Éric de la Clergerie +1
Text simplification aims at making a text easier to read and understand by simplifying grammar and structure while keeping the underlying information identical. It is often conside…