5 papers
RALS: Resources and Baselines for Romanian Automatic Lexical Simplification
Fabian Anghel, Petru Theodor Cristea, Claudiu Creanga +1
We introduce the first dataset that jointly covers both lexical complexity prediction (LCP) annotations and lexical simplification (LS) for Romanian, along with a comparison of lex…
Automatic Correction of Writing Anomalies in Hausa Texts
Ahmad Mustapha Wali, Sergiu Nisioi
Hausa texts are often characterized by writing anomalies, such as incorrect character substitutions and spacing errors, which sometimes hinder natural language processing (NLP) app…
SG-UniBuc-NLP at SemEval-2026 Task 6: Multi-Head RoBERTa with Chunking for Long-Context Evasion Detection
Gabriel Stefan, Sergiu Nisioi
We describe our system for SemEval-2026 Task 6 (CLARITY: Unmasking Political Question Evasions), which classifies English political interview responses by coarse-grained clarity (3…
A Typology of Synthetic Datasets for Dialogue Processing in Clinical Contexts
Steven Bedrick, A. Seza DoÄruöz, Sergiu Nisioi
Synthetic data sets are used across linguistic domains and NLP tasks, particularly in scenarios where authentic data is limited (or even non-existent). One such domain is that of c…
Dialectal and Low-Resource Machine Translation for Aromanian
Alexandru-Iulius Jerpelea, Alina RÄdoi, Sergiu Nisioi
This paper presents the process of building a neural machine translation system with support for English, Romanian, and Aromanian - an endangered Eastern Romance language. The prim…