8 papers
Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction
Mikel Zubillaga, Oscar Sainz, Oier Lopez de Lacalle +1
Document-level Information Extraction (DocIE) aims to produce an output template with the entities, relations, and events of interest occurring in the given document. Standard prac…
Merge and Conquer: Instructing Multilingual Models by Adding Target Language Weights
Eneko Valero, Maria Ribalta i Albado, Oscar Sainz +2
Large Language Models (LLMs) remain heavily centered on English, with limited performance in low-resource languages. Existing adaptation approaches, such as continual pre-training,…
SemBench: A Universal Semantic Framework for LLM Evaluation
Mikel Zubillaga, Naiara Perez, Oscar Sainz +1
Recent progress in Natural Language Processing (NLP) has been driven by the emergence of Large Language Models (LLMs), which exhibit remarkable generative and reasoning capabilitie…
Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque
Oscar Sainz, Naiara Perez, Julen Etxaniz +9
Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages. In this paper, we explore alternatives to…
GuideX: Guided Synthetic Data Generation for Zero-Shot Information Extraction
Neil De La Fuente, Oscar Sainz, Iker GarcÃa-Ferrero +1
Information Extraction (IE) systems are traditionally domain-specific, requiring costly adaptation that involves expert schema design, data annotation, and model training. While La…
Latxa: An Open Language Model and Evaluation Suite for Basque
Julen Etxaniz, Oscar Sainz, Naiara Perez +6
We introduce Latxa, a family of large language models for Basque ranging from 7 to 70 billion parameters. Latxa is based on Llama 2, which we continue pretraining on a new Basque c…