5 papers
SemBench: A Universal Semantic Framework for LLM Evaluation
Mikel Zubillaga, Naiara Perez, Oscar Sainz +1
Recent progress in Natural Language Processing (NLP) has been driven by the emergence of Large Language Models (LLMs), which exhibit remarkable generative and reasoning capabilitie…
Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction
Mikel Zubillaga, Oscar Sainz, Oier Lopez de Lacalle +1
Document-level Information Extraction (DocIE) aims to produce an output template with the entities, relations, and events of interest occurring in the given document. Standard prac…
BERnaT: Basque Encoders for Representing Natural Textual Diversity
Ekhi Azurmendi, Joseba Fernandez de Landa, Jaione Bengoetxea +5
Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally exclude non-standard linguistic varieties, reduce model robus…
HiTZ at VarDial 2025 NorSID: Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation
Jaione Bengoetxea, Mikel Zubillaga, Ekhi Azurmendi +4
In this paper we present our submission for the NorSID Shared Task as part of the 2025 VarDial Workshop (Scherrer et al., 2025), consisting of three tasks: Intent Detection, Slot F…
Event Extraction in Basque: Typologically motivated Cross-Lingual Transfer-Learning Analysis
Mikel Zubillaga, Oscar Sainz, Ainara Estarrona +2
Cross-lingual transfer-learning is widely used in Event Extraction for low-resource languages and involves a Multilingual Language Model that is trained in a source language and ap…