7 papers
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde +2
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimize…
The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations
Carlos Arriaga, Gonzalo MartÃnez, Eneko Sendin +2
The evaluation of large language models is a complex task, in which several approaches have been proposed. The most common is the use of automated benchmarks in which LLMs have to…
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
MarÃa Grandury, Javier Aula-Blasco, Júlia Falcão +22
Leaderboards showcase the current capabilities and limitations of Large Language Models (LLMs). To motivate the development of LLMs that represent the linguistic and cultural diver…
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
Raquel Ferrando, Javier Conde, Gonzalo MartÃnez +1
The computational and energy costs of Large Language Models (LLMs) have increased exponentially driven by the growing model sizes and the massive adoption of LLMs by hundreds of mi…
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
Javier Conde, Miguel González, MarÃa Grandury +3
The evaluation of LLMs has so far focused primarily on how well they can perform different tasks such as reasoning, question-answering, paraphrasing, or translating. For most of th…
Understanding the Impact of Artificial Intelligence in Academic Writing: Metadata to the Rescue
Javier Conde, Pedro Reviriego, JoaquÃn Salvachúa +3
This column advocates for including artificial intelligence (AI)-specific metadata on those academic papers that are written with the help of AI in an attempt to analyze the use of…