activity
20242026
collaborators

5 papers

cs.CL2026

RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian

Andrei-Marius Avram, Aureliu Valentin Antonie, Cosmin-Mircea Croitoru +2

We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews acros…

cs.CL2025

MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language

Andrei-Marius Avram, Ema-Ioana Bănescu, Anda-Teodora Robea +2

This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced be…

cs.CL2025

UniBERT: Adversarial Training for Language-Universal Representations

Andrei-Marius Avram, Marian Lupaşcu, Dumitru-Clementin Cercel +2

This paper presents UniBERT, a compact multilingual language model that uses an innovative training framework that integrates three components: masked language modeling, adversaria…

cs.CL2024

RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation

Andrei-Marius Avram, Mircea Timpuriu, Andreea Iuga +6

Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language proces…

cs.CL2024

RELATE: A Modern Processing Platform for Romanian Language

Vasile Păiş, Radu Ion, Andrei-Marius Avram +2

This paper presents the design and evolution of the RELATE platform. It provides a high-performance environment for natural language processing activities, specially constructed fo…