3 citations · 16 across the 41 of their papers we have counts for
37 papers · 1 filter
ROMPAR: Morphological Completion and Demographic Unlearning for Romanian-Accented Speech Recognition
Andrei-Marius Avram, Aureliu-Valentin Antonie, Ştefan-Bogdan Badea +3
Automated transcription of parliamentary proceedings faces significant hurdles due to demographic bias, dialectal variation, and technical artifacts such as utterance truncation du…
RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian
Andrei-Marius Avram, Aureliu Valentin Antonie, Cosmin-Mircea Croitoru +2
We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews acros…
RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
Mircea Timpuriu, Mihaela-Claudia Cercel, Dumitru-Clementin Cercel
The importance of clear and correct text in legal documents cannot be understated, and, consequently, a grammatical error correction tool meant to assist a professional in the law…
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
George-Andrei Dima, Răzvan-Alexandru Smădu, Dumitru-Clementin Cercel
Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We…
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
Răzvan-Alexandru Smădu, Andreea Iuga, Dumitru-Clementin Cercel +1
Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin…
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
Andrei-Marius Avram, Ema-Ioana Bănescu, Anda-Teodora Robea +2
This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced be…