3 papers
cs.CL2020
The birth of Romanian BERT
Stefan Daniel Dumitrescu, Andrei-Marius Avram, Sampo Pyysalo
Large-scale pretrained language models have become ubiquitous in Natural Language Processing. However, most of these models are available either in high-resource languages, in part…
cs.CL2019
Introducing RONEC -- the Romanian Named Entity Corpus
Stefan Daniel Dumitrescu, Andrei-Marius Avram
We present RONEC - the Named Entity Corpus for the Romanian language. The corpus contains over 26000 entities in ~5000 annotated sentences, belonging to 16 distinct classes. The se…
cs.CL2018
Tools and resources for Romanian text-to-speech and speech-to-text applications
Tiberiu Boros, Stefan Daniel Dumitrescu, Vasile Pais
In this paper we introduce a set of resources and tools aimed at providing support for natural language processing, text-to-speech synthesis and speech recognition for Romanian. Wh…