activity
20192026
most citedAutomatic Spanish Translation of the SQuAD Dataset for Multilingual Question Answering

42 citations · 77 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

TalentCLEF at CLEF2026: Skill and Job Title Intelligence for Human Capital Management

Luis Gasco, Hermenegildo Fabregat, Laura García-Sardiña +5

This paper presents the second edition of the TalentCLEF Challenge, which will run as an evaluation lab as part of CLEF 2026. The aim of TalentCLEF is to promote the development of…

cs.CL2026

JobResQA: A Benchmark for LLM Machine Reading Comprehension on Multilingual Résumés and JDs

Casimiro Pio Carrino, Paula Estrella, Rabih Zbib +2

We introduce JobResQA, a multilingual Question Answering benchmark for evaluating Machine Reading Comprehension (MRC) capabilities of LLMs on HR-specific tasks involving résumés an…

cs.CL2024

MELO: An Evaluation Benchmark for Multilingual Entity Linking of Occupations

Federico Retyk, Luis Gasco, Casimiro Pio Carrino +2

We present the Multilingual Entity Linking of Occupations (MELO) Benchmark, a new collection of 48 datasets for evaluating the linking of entity mentions in 21 languages to the ESC…

cs.CL2023

Promoting Generalized Cross-lingual Question Answering in Few-resource Scenarios via Self-knowledge Distillation

Casimiro Pio Carrino, Carlos Escolano, José A. R. Fonollosa

Despite substantial progress in multilingual extractive Question Answering (QA), models with high and uniformly distributed performance across languages remain challenging, especia…

cs.CL2021★ 23 cited

Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario

Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño +4

This work presents biomedical and clinical language models for Spanish by experimenting with different pretraining choices, such as masking at word and subword level, varying the v…

cs.CL2021★ 4 cited

Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models

Casimiro Pio Carrino, Jordi Armengol-Estapé, Ona de Gibert Bonet +4

We introduce CoWeSe (the Corpus Web Salud Español), the largest Spanish biomedical corpus to date, consisting of 4.5GB (about 750M tokens) of clean plain text. CoWeSe is the result…