42 citations · 77 across the 10 of their papers we have counts for
9 papers · 1 filter
TalentCLEF at CLEF2026: Skill and Job Title Intelligence for Human Capital Management
Luis Gasco, Hermenegildo Fabregat, Laura García-Sardiña +5
This paper presents the second edition of the TalentCLEF Challenge, which will run as an evaluation lab as part of CLEF 2026. The aim of TalentCLEF is to promote the development of…
JobResQA: A Benchmark for LLM Machine Reading Comprehension on Multilingual Résumés and JDs
Casimiro Pio Carrino, Paula Estrella, Rabih Zbib +2
We introduce JobResQA, a multilingual Question Answering benchmark for evaluating Machine Reading Comprehension (MRC) capabilities of LLMs on HR-specific tasks involving résumés an…
MELO: An Evaluation Benchmark for Multilingual Entity Linking of Occupations
Federico Retyk, Luis Gasco, Casimiro Pio Carrino +2
We present the Multilingual Entity Linking of Occupations (MELO) Benchmark, a new collection of 48 datasets for evaluating the linking of entity mentions in 21 languages to the ESC…
Promoting Generalized Cross-lingual Question Answering in Few-resource Scenarios via Self-knowledge Distillation
Casimiro Pio Carrino, Carlos Escolano, José A. R. Fonollosa
Despite substantial progress in multilingual extractive Question Answering (QA), models with high and uniformly distributed performance across languages remain challenging, especia…
Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario
Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño +4
This work presents biomedical and clinical language models for Spanish by experimenting with different pretraining choices, such as masking at word and subword level, varying the v…
Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models
Casimiro Pio Carrino, Jordi Armengol-Estapé, Ona de Gibert Bonet +4
We introduce CoWeSe (the Corpus Web Salud Español), the largest Spanish biomedical corpus to date, consisting of 4.5GB (about 750M tokens) of clean plain text. CoWeSe is the result…