activity
20182024
collaborators

5 papers

cs.CL2024

EuSQuAD: Automatically Translated and Aligned SQuAD2.0 for Basque

Aitor García-Pablos, Naiara Perez, Montse Cuadros +1

The widespread availability of Question Answering (QA) datasets in English has greatly facilitated the advancement of the Natural Language Processing (NLP) field. However, the scar…

cs.CL2020

NUBES: A Corpus of Negation and Uncertainty in Spanish Clinical Texts

Salvador Lima, Naiara Perez, Montse Cuadros +1

This paper introduces the first version of the NUBes corpus (Negation and Uncertainty annotations in Biomedical texts in Spanish). The corpus is part of an on-going research and cu…

cs.CL2020

Sensitive Data Detection and Classification in Spanish Clinical Text: Experiments with BERT

Aitor García-Pablos, Naiara Perez, Montse Cuadros

Massive digital data processing provides a wide range of opportunities and benefits, but at the cost of endangering personal data privacy. Anonymisation consists in removing or rep…

cs.CL2018

Hate Speech Dataset from a White Supremacy Forum

Ona de Gibert, Naiara Perez, Aitor García-Pablos +1

Hate speech is commonly defined as any communication that disparages a target group of people based on some characteristic such as race, colour, ethnicity, gender, sexual orientati…

cs.CL2018

Biomedical term normalization of EHRs with UMLS

Naiara Perez, Montse Cuadros, German Rigau

This paper presents a novel prototype for biomedical term normalization of electronic health record excerpts with the Unified Medical Language System (UMLS) Metathesaurus. Despite…