75 citations · 78 across the 2 of their papers we have counts for
4 papers
From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern French
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz +4
Language models for historical states of language are becoming increasingly important to allow the optimal digitisation and analysis of old textual sources. Because these historica…
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
Pedro Javier Ortiz Suárez, Laurent Romary, Benoît Sagot
We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) fo…
Establishing a New State-of-the-Art for French Named Entity Recognition
Pedro Javier Ortiz Suárez, Yoann Dupont, Benjamin Muller +2
The French TreeBank developed at the University Paris 7 is the main source of morphosyntactic and syntactic annotations for French. However, it does not include explicit informatio…
CamemBERT: a Tasty French Language Model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez +5
Pretrained language models are now ubiquitous in Natural Language Processing. Despite their success, most available models have either been trained on English data or on the concat…