2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.CL2022
From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern French
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz +4
Language models for historical states of language are becoming increasingly important to allow the optimal digitisation and analysis of old textual sources. Because these historica…
cs.CL2020★ 2 cited
Standardizing linguistic data: method and tools for annotating (pre-orthographic) French
Simon Gabay, Thibault Clérice, Jean-Baptiste Camps +2
With the development of big corpora of various periods, it becomes crucial to standardise linguistic annotation (e.g. lemmas, POS tags, morphological annotation) to increase the in…