9 citations · 10 across the 5 of their papers we have counts for
5 papers
Data Contamination Report from the 2024 CONDA Shared Task
Oscar Sainz, Iker García-Ferrero, Alon Jacovi +25
The 1st Workshop on Data Contamination (CONDA 2024) focuses on all relevant aspects of data contamination in natural language processing, where data contamination is understood as…
Event Extraction in Basque: Typologically motivated Cross-Lingual Transfer-Learning Analysis
Mikel Zubillaga, Oscar Sainz, Ainara Estarrona +2
Cross-lingual transfer-learning is widely used in Event Extraction for low-resource languages and involves a Multilingual Language Model that is trained in a source language and ap…
NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Oscar Sainz, Jon Ander Campos, Iker García-Ferrero +3
In this position paper, we argue that the classical evaluation on Natural Language Processing (NLP) tasks using annotated benchmarks is in trouble. The worst kind of data contamina…
IXA/Cogcomp at SemEval-2023 Task 2: Context-enriched Multilingual Named Entity Recognition using Knowledge Bases
Iker García-Ferrero, Jon Ander Campos, Oscar Sainz +2
Named Entity Recognition (NER) is a core natural language processing task in which pre-trained language models have shown remarkable performance. However, standard benchmarks like…
What do Language Models know about word senses? Zero-Shot WSD with Language Models and Domain Inventories
Oscar Sainz, Oier Lopez de Lacalle, Eneko Agirre +1
Language Models are the core for almost any Natural Language Processing system nowadays. One of their particularities is their contextualized representations, a game changer featur…