5 papers
RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian
Andrei-Marius Avram, Aureliu Valentin Antonie, Cosmin-Mircea Croitoru +2
We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews acros…
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
Andrei-Marius Avram, Ema-Ioana BÄnescu, Anda-Teodora Robea +2
This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced be…
UniBERT: Adversarial Training for Language-Universal Representations
Andrei-Marius Avram, Marian LupaÅcu, Dumitru-Clementin Cercel +2
This paper presents UniBERT, a compact multilingual language model that uses an innovative training framework that integrates three components: masked language modeling, adversaria…
RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation
Andrei-Marius Avram, Mircea Timpuriu, Andreea Iuga +6
Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language proces…
RELATE: A Modern Processing Platform for Romanian Language
Vasile PÄiÅ, Radu Ion, Andrei-Marius Avram +2
This paper presents the design and evolution of the RELATE platform. It provides a high-performance environment for natural language processing activities, specially constructed fo…