collaborators

5 papers

cs.CL2026

RALS: Resources and Baselines for Romanian Automatic Lexical Simplification

Fabian Anghel, Petru Theodor Cristea, Claudiu Creanga +1

We introduce the first dataset that jointly covers both lexical complexity prediction (LCP) annotations and lexical simplification (LS) for Romanian, along with a comparison of lex…

cs.CL2026

Automatic Correction of Writing Anomalies in Hausa Texts

Ahmad Mustapha Wali, Sergiu Nisioi

Hausa texts are often characterized by writing anomalies, such as incorrect character substitutions and spacing errors, which sometimes hinder natural language processing (NLP) app…

cs.CL2026

SG-UniBuc-NLP at SemEval-2026 Task 6: Multi-Head RoBERTa with Chunking for Long-Context Evasion Detection

Gabriel Stefan, Sergiu Nisioi

We describe our system for SemEval-2026 Task 6 (CLARITY: Unmasking Political Question Evasions), which classifies English political interview responses by coarse-grained clarity (3…

cs.CL2026

A Typology of Synthetic Datasets for Dialogue Processing in Clinical Contexts

Steven Bedrick, A. Seza Doğruöz, Sergiu Nisioi

Synthetic data sets are used across linguistic domains and NLP tasks, particularly in scenarios where authentic data is limited (or even non-existent). One such domain is that of c…

cs.CL2025

Dialectal and Low-Resource Machine Translation for Aromanian

Alexandru-Iulius Jerpelea, Alina Rădoi, Sergiu Nisioi

This paper presents the process of building a neural machine translation system with support for English, Romanian, and Aromanian - an endangered Eastern Romance language. The prim…