activity
20232026
most citedNLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

9 citations · 9 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CL2026

Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction

Mikel Zubillaga, Oscar Sainz, Oier Lopez de Lacalle +1

Document-level Information Extraction (DocIE) aims to produce an output template with the entities, relations, and events of interest occurring in the given document. Standard prac…

cs.CL2025

Automatic Essay Scoring and Feedback Generation in Basque Language Learning

Ekhi Azurmendi, Xabier Arregi, Oier Lopez de Lacalle

This paper introduces the first publicly available dataset for Automatic Essay Scoring (AES) and feedback generation in Basque, targeting the CEFR C1 proficiency level. The dataset…

cs.AI2025

REMoH: A Reflective Evolution of Multi-objective Heuristics approach via Large Language Models

Diego Forniés-Tabuenca, Alejandro Uribe, Urtzi Otamendi +3

Multi-objective optimization is fundamental in complex decision-making tasks. Traditional algorithms, while effective, often demand extensive problem-specific modeling and struggle…

cs.CL2025

Vision-Language Models Struggle to Align Entities across Modalities

Iñigo Alonso, Gorka Azkune, Ander Salaberria +2

Cross-modal entity linking refers to the ability to align entities and their attributes across different modalities. While cross-modal entity linking is a fundamental skill needed…

cs.CL2024

Event Extraction in Basque: Typologically motivated Cross-Lingual Transfer-Learning Analysis

Mikel Zubillaga, Oscar Sainz, Ainara Estarrona +2

Cross-lingual transfer-learning is widely used in Event Extraction for low-resource languages and involves a Multilingual Language Model that is trained in a source language and ap…

cs.CL20239 cited

NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Oscar Sainz, Jon Ander Campos, Iker García-Ferrero +3

In this position paper, we argue that the classical evaluation on Natural Language Processing (NLP) tasks using annotated benchmarks is in trouble. The worst kind of data contamina…