activity
20202026
most citedSkinDistilViT: Lightweight Vision Transformer for Skin Lesion Classification

3 citations · 16 across the 41 of their papers we have counts for

collaborators
Showing cs.CLShow all

37 papers · 1 filter

cs.CL2026

ROMPAR: Morphological Completion and Demographic Unlearning for Romanian-Accented Speech Recognition

Andrei-Marius Avram, Aureliu-Valentin Antonie, Ştefan-Bogdan Badea +3

Automated transcription of parliamentary proceedings faces significant hurdles due to demographic bias, dialectal variation, and technical artifacts such as utterance truncation du…

cs.CL2026

RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian

Andrei-Marius Avram, Aureliu Valentin Antonie, Cosmin-Mircea Croitoru +2

We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews acros…

cs.CL2026

RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian

Mircea Timpuriu, Mihaela-Claudia Cercel, Dumitru-Clementin Cercel

The importance of clear and correct text in legal documents cannot be understated, and, consequently, a grammatical error correction tool meant to assist a professional in the law…

cs.CL2025

Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models

George-Andrei Dima, Răzvan-Alexandru Smădu, Dumitru-Clementin Cercel

Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We…

cs.CL20251 cited

SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset

Răzvan-Alexandru Smădu, Andreea Iuga, Dumitru-Clementin Cercel +1

Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin…

cs.CL2025

MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language

Andrei-Marius Avram, Ema-Ioana Bănescu, Anda-Teodora Robea +2

This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced be…