8 papers
RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
Mircea Timpuriu, Mihaela-Claudia Cercel, Dumitru-Clementin Cercel
The importance of clear and correct text in legal documents cannot be understated, and, consequently, a grammatical error correction tool meant to assist a professional in the law…
RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams
Andrei Vlad Man, RÄzvan-Alexandru SmÄdu, Cristian-George Craciun +3
The intersection of AI and legal systems presents a growing need for tools that support legal education, particularly in under-resourced languages such as Romanian. In this work, w…
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
Andrei-Marius Avram, Ema-Ioana BÄnescu, Anda-Teodora Robea +2
This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced be…
SaRoHead: Detecting Satire in a Multi-Domain Romanian News Headline Dataset
Mihnea-Alexandru Vîrlan, RÄzvan-Alexandru SmÄdu, Dumitru-Clementin Cercel +2
The primary goal of a news headline is to summarize an event in as few words as possible. Depending on the media outlet, a headline can serve as a means to objectively deliver a su…
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering
Cristian-George CrÄciun, RÄzvan-Alexandru SmÄdu, Dumitru-Clementin Cercel +1
Pre-trained Language Models (PLMs) have shown remarkable performances in recent years, setting a new paradigm for NLP research and industry. The legal domain has received some atte…
RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation
Andrei-Marius Avram, Mircea Timpuriu, Andreea Iuga +6
Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language proces…