3 papers
cs.CL2020
The birth of Romanian BERT
Stefan Daniel Dumitrescu, Andrei-Marius Avram, Sampo Pyysalo
Large-scale pretrained language models have become ubiquitous in Natural Language Processing. However, most of these models are available either in high-resource languages, in part…
cs.CL2020
UPB at SemEval-2020 Task 6: Pretrained Language Models for Definition Extraction
Andrei-Marius Avram, Dumitru-Clementin Cercel, Costin-Gabriel Chiru
This work presents our contribution in the context of the 6th task of SemEval-2020: Extracting Definitions from Free Text in Textbooks (DeftEval). This competition consists of thre…
cs.CL2019
Introducing RONEC -- the Romanian Named Entity Corpus
Stefan Daniel Dumitrescu, Andrei-Marius Avram
We present RONEC - the Named Entity Corpus for the Romanian language. The corpus contains over 26000 entities in ~5000 annotated sentences, belonging to 16 distinct classes. The se…