collaborators

7 papers

cs.CL2026

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde +2

Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimize…

cs.AI2025

The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations

Carlos Arriaga, Gonzalo Martínez, Eneko Sendin +2

The evaluation of large language models is a complex task, in which several approaches have been proposed. The most common is the use of automated benchmarks in which LLMs have to…

cs.CL2025

La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America

María Grandury, Javier Aula-Blasco, Júlia Falcão +22

Leaderboards showcase the current capabilities and limitations of Large Language Models (LLMs). To motivate the development of LLMs that represent the linguistic and cultural diver…

cs.CL2025

Is There a Case for Conversation Optimized Tokenizers in Large Language Models?

Raquel Ferrando, Javier Conde, Gonzalo Martínez +1

The computational and energy costs of Large Language Models (LLMs) have increased exponentially driven by the growing model sizes and the massive adoption of LLMs by hundreds of mi…

cs.CL2025

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans

Javier Conde, Miguel González, María Grandury +3

The evaluation of LLMs has so far focused primarily on how well they can perform different tasks such as reasoning, question-answering, paraphrasing, or translating. For most of th…

cs.AI2025

Understanding the Impact of Artificial Intelligence in Academic Writing: Metadata to the Rescue

Javier Conde, Pedro Reviriego, Joaquín Salvachúa +3

This column advocates for including artificial intelligence (AI)-specific metadata on those academic papers that are written with the help of AI in an attempt to analyze the use of…