3 papers
cs.CL2025
MessIRve: A Large-Scale Spanish Information Retrieval Dataset
Francisco Valentini, Viviana Cotik, Damián Furman +3
Information retrieval (IR) is the task of finding relevant documents in response to a user query. Although Spanish is the second most spoken native language, there are few Spanish…
cs.CL2025
Indigenous Languages Spoken in Argentina: A Survey of NLP and Speech Resources
Belu Ticona, Fernando Carranza, Viviana Cotik
Argentina has a large yet little-known Indigenous linguistic diversity, encompassing at least 40 different languages. The majority of these languages are at risk of disappearing, r…
cs.CL2024
Exploring Large Language Models for Hate Speech Detection in Rioplatense Spanish
Juan Manuel Pérez, Paula Miguel, Viviana Cotik
Hate speech detection deals with many language variants, slang, slurs, expression modalities, and cultural nuances. This outlines the importance of working with specific corpora, w…