4 papers
Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English
Jeremy Barnes, Naiara Perez, Alba Bonet-Jover +1
Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models. However, these…
Societal Alignment Frameworks Can Improve LLM Alignment
Karolina Stańczak, Nicholas Meade, Mehar Bhatia +14
Recent progress in large language models (LLMs) has focused on producing responses that meet human expectations and align with shared values - a process coined alignment. However,…
Conditioning LLMs to Generate Code-Switched Text
Maite Heredia, Gorka Labaka, Jeremy Barnes +1
Code-switching (CS) is still a critical challenge in Natural Language Processing (NLP), due to the limited availability of large-scale, diverse CS datasets for robust training and…
Truth Knows No Language: Evaluating Truthfulness Beyond English
Blanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz +4
We introduce a professionally translated extension of the TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. Truthfulness evaluations…